Smarsh vs Theta Lake: how they compare in 2026
Smarsh and Theta Lake capture, archive and supervise the communications of regulated firms, from email and chat to voice, video and mobile messaging, and both run AI over that record to flag risk. Theta Lake sits in the top two bands on fourteen of fifteen axes and Smarsh on ten of fifteen, with identical grades on ten and Theta Lake higher on the rest. The difference is how each describes its AI. Theta Lake holds ISO 42001 certification for AI management, links every detection to the captured interaction behind it, and states in its agreement that output is machine generated, may be wrong and must be reviewed by the customer. Its data processing addendum commits to deletion on request and gives a right to object to new subprocessors. Smarsh's Noise Reduction Agent suppresses surveillance alerts, and nothing published says whether suppressed alerts can be reviewed or sampled, a live question where supervision is a regulatory duty. Its terms do not say whether client communications train its models. Smarsh answers with audit paperwork, committing by contract to hand clients its ISO 27001 and SSAE 18 reports and a penetration test summary.
At a glance
All 15 axes, side by side
The same grid applied to every vendor in the index, graded from public sources. Hover a grade to see what the letter means on that axis.
AI Centrality
How much of the product is actually AI. Whether the machine learning is the mechanism the buyer is paying for or a feature layered onto conventional software, and whether the vendor is specific about which is which.
The models drive the capabilities the vendor now leads with, on a capture and archive platform that works without them, which is the B band. The Intelligent Agent and Noise Reduction Agent filter and detect risk in surveillance, the AI Assistant summarises and translates, and the Discovery Agent produces summaries, timelines and custodian mapping for investigations. Underneath is a capture, retention, legal hold, search and export platform, sold for decades as a recordkeeping system of record under SEC Rule 17a-4, that functions fully without models. Verified 18 September 2026.
Remove the models and a capture, archive, search and legal hold platform remains, which is the B band exactly. Unified Capture and Unified Search and Archiving are ingestion and retrieval across more than 100 certified API integrations, and the SEC 17a-4 WORM archive, the eDiscovery export and the custodian legal hold workflow all function without a model. Proactive Compliance is where the models are the engine of a core capability: patented machine learning and NLP detection across what is shared, shown, spoken and typed, more than 80 built-in and custom policies driving that detection, and the aiComms classifiers for prompt injection, jailbreak behaviour, shadow AI and unethical summary steering. The vendor's own framing on the AI Communication and Interaction Governance page, modified 14 July 2026, is a governance and detection layer added over the capture substrate. Verified 12 September 2026.
Citation Accuracy and Hallucination Disclosure
Whether the vendor publishes measured accuracy on citations and assertions, grounds output to primary sources, and says plainly what its system does when it does not know. Legal has a documented public record of fabricated citations reaching filed briefs, so an untested claim of accuracy is not evidence.
Performance is asserted with figures that carry no method, and grounding is claimed rather than described, which is the C band. The vendor's releases and pages say the Noise Reduction Agent cuts false positives by 60 per cent, review volumes fall by up to 50 per cent, the Intelligent Agent surfaces three to five times more real risk, and the Discovery Agent cuts investigation costs by up to 75 per cent, with no sample, baseline or test described. The innovations FAQ says AI summaries and risk signals rest on traceable source data and are reproducible, but does not describe how a summary links back to the messages it relies on or what happens when the AI misreads a conversation. Verified 18 September 2026.
R15 governs. This product cites no legal authority, so the citator limb and the primary-authority limb do not apply to the product class and are neither credited nor penalised. What bites is grounding and hallucination disclosure, and both are real. Every detection resolves to the captured interaction it came from, with contextual investigation views, timeline views, replay and add-to-case navigation documented on the product and AI governance pages, so a reviewer opens the source content rather than a summary of it, and Proactive Compliance publishes built-in audit and explainability reporting for ML and AI systems. The hallucination disclosure sits in the agreement rather than the marketing: MSA clause 10(d) states that output is generated by machine learning capabilities, warrants nothing as to accuracy, completeness or reliability, notes that output may differ between runs, and places evaluation on the customer including human review. What keeps this off A is that no measured detection accuracy, recall or precision figure is published on any surface located, and no test set is described. MSA updated 12 March 2026; verified 12 September 2026.
Autonomy and Oversight Model
What the system decides on its own, what a lawyer must approve, and whether the vendor documents where the review point sits. A tool that drafts under review and a tool that files without one are different products and different risks.
Automated suppression of surveillance alerts is described without a published human control, which places this at C. The Noise Reduction Agent and Intelligent Agent suppress low-relevance alerts so supervisory teams see fewer, and the vendor describes these as autonomous systems that augment rather than replace human expertise, with audit trails and chain of custody. Nothing published says whether suppressed alerts can be reviewed or are sampled, what threshold governs suppression, or who approves the agent's configuration, which matters because supervisory review of communications is a regulatory duty for the vendor's financial services customers. Discovery Agent output is presented as input to investigators. Verified 18 September 2026.
Review surfaces are real and documented: automated multi-party review workflows that route by detection, contextual investigation and timeline views, role-based access control on private links into those views, bidirectional alert integration with SIEM and SOC tools, and customer-parametrised compound detection rules that set what fires. The limb the B band names as commonly absent is absent here. The vendor states in the same material that the system acts alone in places -- patented risk remediation and prevention, automatic application of legal hold to AI summaries for custodians in legal matters, real-time policy notifications and disclaimers inserted into Microsoft Teams conversations, and dynamic retention rules deciding what is not retained -- and the threshold at which any of that proceeds without a reviewer is not published. R37 rule 2 governs: the tension between assisting reviewers and acting automatically is not a reason to downgrade for the contradiction as such, it identifies the missing limb, and the grade goes there. MSA clause 10(d) allocates evaluation of output to the customer. Verified 12 September 2026.
Operational and Outcome Evidence
Named, dated evidence that the product works in production at real firms or legal departments. Case studies with figures and identified customers count. Unattributed testimonials and launch announcements do not.
A commissioned study with a described method and a named customer with figures, neither both at once, which is the B band. Smarsh publishes a Forrester Consulting Total Economic Impact study (September 2024) that interviewed six representatives of financial services firms, built a composite organisation, and risk-adjusted each benefit: 124 per cent ROI over three years, archive costs down 25 per cent, e-discovery time down 65 per cent, and false-positive surveillance alerts down 25 to 45 per cent by interviewee estimate; the interviewees are anonymous and Smarsh chose them. Its customer story for Securities America, a named broker-dealer, reports messages flagged for review falling from 19 to 10 per cent (8,000 a day) in three months, but that 2020 result came from professional services tuning of policy rules, not the current AI agents, and no method is given. The 2026 release figures for the AI agents still carry no method. Verified 18 September 2026.
A named customer with a named role and a specific deployment: the Head of Technology at Longview Partners, on MiFID II compliance for Microsoft Teams, quoted on the vendor's own site. No figures accompany it, which is the B band's stated shape of a named customer without measurement. Analyst recognition is extensive and is recorded rather than credited, because analyst placement is not deployment evidence: Furthest in Vision in the 2025 Gartner Magic Quadrant for Digital Communications Governance and Archiving, ranked first in five of six use cases in the Critical Capabilities companion, and a Gartner Peer Insights listing. A case studies library is published and was not opened. Verified 12 September 2026.
Privilege and Confidentiality Posture
How client confidences are handled: attorney client privilege and work product treatment, segregation of one client matter from another, whether client data trains any model, and what the vendor commits to in writing rather than in marketing.
Contractual confidentiality is in place while the AI data position is not addressed, which places this at C. The Smarsh Services Agreement treats the client's data as the client's property and confidential information, with notice before compelled disclosure, and licenses Smarsh to use client data to provide support and improve the services on the client's behalf. No published term addresses whether client communications are used to train or adapt the domain-adapted models behind the AI agents, which model providers if any see client data, or privilege and work product in data held for discovery. The innovations FAQ describes AI running inside governed environments with role-based access and chain of custody. Verified 18 September 2026.
The commitments are contractual and readable before signing, which is more than most of this corpus offers, and the A band's privilege limb is absent, which R33 makes decisive. Published and read: MSA section 1(d) defines Customer Data, Reports and Output as the customer's Confidential Information; the data licence in section 5(a) confines Theta Lake's use of Customer Data to providing the Service, generating Output and Reports, and creating Usage and Anonymized Data, for the Subscription Period only; the DPA's United States schedule states the vendor will not retain, use, disclose, sell or share Personal Data other than to provide the Services on documented instructions, and will not combine it with data from other entities; DPA section 8 commits to deletion or return on written request; retention is customer-set; and MSA section 8(c) commits to notice before any compelled disclosure. Two limbs keep this at B: nothing located addresses privilege or work product treatment, and no position is published on what any third-party model provider may retain. MSA updated 12 March 2026, DPA updated 2 July 2025, both read in full on the support portal; verified 12 September 2026.
UPL and Professional Responsibility Posture
Whether the vendor is clear that it supplies a tool rather than legal advice, who its audience is, and how it addresses unauthorized practice of law, competence and supervision duties, and jurisdiction limits. ABA Formal Opinion 512 is the reference point. Where the advice line is not the duty a product raises, the axis is read through the nearest professional duty it does raise: judicial conduct rules and the reviewing duty for products sold only to courts, and the duty to bill for time actually spent for products that draft time entries.
A clear position that the product is a compliance tool and not legal advice, short of the supervision dimension, which is the B band. The Services Agreement states that Smarsh does not guarantee that use of the services, or its advice or consulting, will ensure the client's legal compliance, and the legal documents page says Smarsh materials are not legal advice and customers must consult an attorney. The audience is compliance, legal and records professionals in regulated firms. Nothing addresses how supervisors or counsel should treat AI-suppressed alerts or AI summaries in meeting their own obligations. Verified 18 September 2026.
A real published position on advice versus tooling, and it sits in the instrument that governs the product rather than in a website notice. MSA clause 10(d) states that output is machine-generated, may contain errors and misstatements, may be incomplete or inaccurate, does not represent Theta Lake's views, and is the customer's sole responsibility to evaluate for accuracy and appropriateness including by human review. Who may use it is stated as well: Authorized Users named under an Order Form, for the customer's internal business operations. Real Time Advisor delivers the customer's own policies, training links and disclaimers rather than legal conclusions. Short of the full band on two counts: nothing addresses a supervising lawyer's competence and supervision duties, and no jurisdiction limits are stated for a product sold across the US, UK and EU. R15 applies to the consumer-facing limb, which does not bite on an enterprise platform with no public-facing advice surface. R50 noted: the separate website terms of use govern the site and are not the instrument graded here. Verified 12 September 2026.
AI Governance and Bias Disclosure
Published governance over model behavior: who owns it inside the vendor, what is tested before release, and what is disclosed about disparate output across matter types, parties, or populations.
Transparency features are described without a governance framework, which places this at C. The innovations FAQ says AI runs within preserved, governed environments with full audit trails, traceable source data, chain of custody and role-based access so outputs are transparent and reproducible, and press coverage quotes the vendor on auditability aligned with regulator expectations. No accountable owner, pre-release testing regime or disclosure of uneven performance across languages or communication styles is published, although the misconduct detection agent is pitched on reading slang, jargon and multilingual exchanges. Smarsh also sells AI governance products to financial firms; under R126 that purpose earns nothing on this row. Verified 18 September 2026.
ISO/IEC 42001 certification, and R36 sets the band: an independently audited AI management standard is real substance short of testing results or a named owner. This record carries more than the Ontra precedent and still reaches none of the A limbs. Published: ISO/IEC 42001:2023, a CSA AI Trustworthy Pledge 2025 listing on the trust centre, and classifiers described as certified, safe and transparent for detecting AI interaction risks. Absent: no accountable owner inside the vendor is named, no pre-release testing regime is described, and nothing is published about uneven output across content types, languages or populations for a detection product whose false negatives are themselves a compliance failure. One discrepancy belongs on the record: the AI governance page claims CSA STAR for AI Level 2 while the trust centre's own compliance list shows CSA STAR Level 1. A 'How We Use AI' security document is named on the trust centre behind an access request and was not obtained. Verified 12 September 2026.
AI Safety and Data Stewardship
Retention, deletion, access control, and what happens to prompts and documents after they are processed. Whether the vendor states its subprocessors and its incident practice, or leaves the buyer to assume.
The agreement covers retention, deletion, access and part of the subprocessor picture, short of a published incident commitment, which is the B band. The Services Agreement lets the client set retention periods (up to seven years by default in Professional Archive), sets temporary retention of up to 30 days for capture services with deletion after, provides for deletion of client data after termination, and lists subprocessors for mobile capture by name and location; the trust page states encryption in transit and at rest, SSO and MFA, and regular penetration testing. The Information Security Addendum and DPA are available on request and were not read, no incident notification timeline is published, and the agreement's post-termination terms conflict between deletion as soon as practicable and retention for up to six months. Verified 18 September 2026.
All five A limbs are published, contractual and specific. Retention is customer-set with explicit non-retention available, recorded in DPA Annex II as limited retention on durations defined by the data exporter and in the privacy policy as customer control with multiple retention periods. Deletion sits in DPA section 8, with deletion certification on written request under the SCC modifications and self-service export to the customer's own S3 bucket, Azure Blob container or O365 account under MSA section 4(b). Access control covers encryption in transit and at rest, customer-specific keys with an option to manage them independently in AWS or Azure KMS, SAML single sign-on, role-based access control and two-factor authentication. Incident practice is a contractual commitment in DPA section 4(b) to notify without undue delay with available detail, mitigation taken and recommended customer steps. The fifth limb, subprocessor disclosure, is satisfied on the agreement's own terms: the DPA obliges the vendor, on engaging any new subprocessor, to update the Subprocessor Site with that subprocessor's name, location and the activities it will perform; a customer may object in writing on reasonable grounds relating to the protection of personal data, the parties must work in good faith toward a resolution, and if none is reached within thirty days the customer may terminate the agreement as its sole and exclusive remedy. Execution of the DPA is also execution of the Standard Contractual Clauses and their annexes. AMENDED 12 September 2026, from B, on evidence located after the row was first written. The row previously sat at B because the subprocessor list itself was refused to this index's fetcher and R25 forbids asserting a top band on an unopened artifact. R25 addresses artifacts that are ungated and were simply not read; this one is machine-refused, and holding the grade down for that would turn a limit on the reader into a finding against the vendor, which section 6.6 forbids. What the vendor publishes is established: a contractual subprocessor disclosure obligation, a named site carrying it, change notification, an objection right and a termination remedy. The identity of the individual subprocessors was not read and is not asserted here. Verified 12 September 2026.
AI Liability and Recourse
What the vendor stands behind contractually when its output is wrong. Indemnities, caps, carve outs, insurance, and whether any of it is published or only reachable through a negotiated agreement.
A detailed published liability position with an indemnity and cap, with nothing on AI output, which is the B band. The Services Agreement gives a vendor indemnity against claims that use of the services infringes a U.S. patent, trademark or copyright, caps Smarsh's liability at fees received in the prior twelve months, provides service-level credits as the remedy for availability failures, and disclaims any guarantee that use of the services ensures legal compliance. No term addresses AI-generated summaries or suppressed alerts, and no insurance position is published. Verified 18 September 2026.
An unusually complete published position that still stops short of what the A band asks. MSA section 11(a) gives a defence and indemnity against third-party patent, trademark and copyright claims arising from the Theta Lake Assets including the customer's permitted use, with named carve-outs for unauthorised use, reproduction or modification, a repair-replace-refund election, and an express statement that this is the sole and exclusive remedy for IP claims. Section 12 caps collective liability at the fees received in the twelve months before the event and defines Uncapped Claims to include gross negligence, recklessness, intentional misconduct and violation of the other party's IP. Exhibit A warrants 99.999% system availability with a service-credit schedule at three thresholds and a claim procedure a customer can actually run. What the vendor stands behind when its output is wrong is answered expressly and in the negative: clause 10(d) disclaims all warranty as to accuracy, completeness and reliability of machine-generated output, no output indemnity is offered, and no insurance is named anywhere located. The A band asks what the vendor stands behind when the system is wrong; the published answer is nothing, which is complete disclosure rather than a top-band position. R50 noted: the website terms of use are a different instrument and do not grade the platform. Verified 12 September 2026.
Practice Systems Integration Depth
How deeply the product reaches into the systems legal work already lives in: document management such as iManage and NetDocuments, Word and Outlook, contract lifecycle management, matter management, e-billing, and court filing systems.
Broad named connections with some depth described, short of documentation read, which is the B band. The Services Agreement and pages describe capture from email, collaboration, mobile carriers such as Verizon and AT&T, apps such as WhatsApp and Signal, voice, social media, websites and Microsoft 365 Copilot, and the platform offers Audit, Identity and Review Alert APIs and exports to outside counsel tools. Product documentation sits in the Smarsh Central support portal and was not read. Verified 18 September 2026.
More than 100 API-based capture integrations, built in-house and certified by the platform partners, published by category with a page for each: Microsoft Teams, Zoom, Webex by Cisco, RingCentral, Slack, Symphony, Verizon, Asana, Mural, CrowdStrike Falcon Next-Gen SIEM and others, across unified communications, contact centre, whiteboards, content management, mobile and text, email, voice, cloud storage and archives, social and financial messaging. Depth is described where it matters for evidence: what is captured per modality, bidirectional alert integration with SIEM and SOC tools, an open developer platform with endpoints to automate workflows and extract interactions and enrichment data, export to the customer's own S3, Azure Blob or O365 under MSA section 4(b), and a documented Relativity integration for full-context eDiscovery and legal hold. The A band asks for the systems legal work already lives in, and that is where the estate thins: Relativity is the one legal system integrated, and no document management or practice management connection was located. Verified 12 September 2026.
Deployment Model and Data Residency
Where the software runs and where the data sits. Multi tenant cloud, single tenant, private deployment, on premises, and whether region of residence is a published option or an enterprise conversation.
Hosting model and region are stated per service in the agreement, without the AI processing location, which is the B band. The Services Agreement states that Professional Archive and Web Archive run in a Smarsh-managed environment in the United States, Cloud Capture runs in a multi-tenant AWS environment in the United States, and mobile capture stores data in the United States unless agreed otherwise. Where the AI agents process data, and whether other regions are offered for them, is not stated in the documents read. Verified 18 September 2026.
Tenancy is stated and residency is offered, which is B on the band as amended. The security architecture page states dedicated server environments at AWS and Azure. The AI governance page states that customers can selectively collect and apply dynamic archiving retention rules deciding which records to keep, for how long and in what region, with the ability to store in any and multiple locations to meet any state or national data sovereignty requirement, and the trust centre carries separate AWS and Azure infrastructure entries. What is not published is the list of regions actually available, what changes between tiers, and where processing happens as distinct from where data is stored, which is the separation the A band asks for. A regional endpoint is visible in a product screenshot. R38 is satisfied on tenancy alone in any event, and the residency material here goes beyond that. Verified 12 September 2026.
Security Certifications and Trust Center
Independent attestation a buyer can pull without a sales call: SOC 2, ISO 27001, penetration test summaries, a trust center with current reports and named scope rather than a badge image.
Named certifications with a contractual route to the reports, short of evidence read, which is the B band. The trust page shows ISO 27001 certification under ANAB accreditation and third-party SOC audits, and the Services Agreement commits Smarsh to annual independent audits under ISO 27001 or SSAE 18 and to give clients its most recent ISO 27001 and SSAE 18 reports and a penetration test summary as standard audit documentation. The certification scope, auditor and report periods are not stated on any page read. The trust.smarsh.com portal, hosted on Vanta, was opened and returned only its page description, so its contents could not be read. Verified 18 September 2026.
A live SafeBase trust centre at trust.thetalake.com, reachable without a sales call and not linked from the main site's navigation or footer, naming SOC 2, PCI DSS, SEC Rule 17a-4, ISO/IEC 42001:2023, ISO/IEC 27001, TruSight, CSA STAR Level 1, the CSA AI Trustworthy Pledge 2025, GDPR, CCPA, CPRA and PIPEDA, and listing the artifacts behind them: SOC 2 report, ISO 42001 certificate, PCI-DSS AOC, penetration test report, Information Security Policy, STAR3 security architecture and a subprocessors entry. The privacy policy states the SOC 2 Type II and PCI DSS audits are annual and that SOC 2 controls are mapped to ISO 27001 and HIPAA, and the MSA makes the most recent audit report available to customers on the support portal. What holds this at B is that the reports sit behind a Get access request whose tier the portal does not state, and no attestation scope or date appears on the ungated surface, so R5's closing rule applies: describe what was seen and grade the lower tier. Two discrepancies belong on the record: the trust centre lists CSA STAR Level 1 while the AI governance page claims CSA STAR for AI Level 2, and the trust centre lists ISO/IEC 27001 as a compliance item while the security architecture page says only that controls are aligned with it. No auditor is named on any ungated surface. Trust centre read 12 September 2026.
Model Supply Chain Disclosure
Which models sit underneath, whose they are, where they run, and whether the vendor commits to telling customers when that changes. A legal buyer inherits every dependency it cannot see.
Models are referred to without being identified, which is the C band. The vendor describes production-ready AI models for global institutions, press coverage quotes it on domain-adapted large language models built with an in-house team, and its APIs support customers' own models; no model, provider or version is named in the agreement or on the pages read, and nothing commits to notice when the models behind surveillance decisions change. Verified 18 September 2026.
The detection stack is described as patented, in-house machine learning and natural language processing with built-in classifiers and compound detection rules, and nothing published names a model or identifies a provider underneath it, which is the C band. The AI vendors named on the integrations page -- Anthropic, OpenAI, Microsoft Copilot, Zoom AI Companion -- are the objects this product governs, not disclosed suppliers to it, and spending a governance fact on a supply-chain axis would be the double-credit error the ground rules name. Customers may bring their own classification tools and models to run against aiComms, which is extensibility on the customer's side rather than disclosure of the vendor's. A 'How We Use AI' security document is named on the trust centre behind an access request and was not obtained: gated, not absent. R34 noted: the DPA's subprocessor change-notification commitment cannot move this axis while the models are unnamed, and it is credited on the outside counsel guideline signal instead. Verified 12 September 2026.
Commercial Transparency
Whether a buyer can learn what this costs without entering a sales process: published rates, the unit being charged, what sits behind an enterprise tier, and what implementation adds.
The unit and structure are published without the figure, which is the B band. The Services Agreement explains that Professional Archive is licensed per connection (a mailbox, account, phone number or social profile) and Web Archive per domain and per page, with minimum commitments equal to the recurring fees, usage-based overage fees, a renewal uplift capped at ten per cent, and additional fees for retention beyond seven years; the Intelligent Agent was announced as a priced add-on. No price or rate is published and buying runs through sales. Regraded from C on 18 September 2026 under R45: the band text for B names unit and structure without the figure. Verified 18 September 2026.
Real pricing is published for part of the charge, which is the B band. The vendor's own AWS Marketplace listing, where Theta Lake is the seller of record and the listing content is the vendor's, publishes two platform tiers with annual figures -- SMB covering up to 999 users at $15,000 for a twelve-month contract, Enterprise covering 1,000 or more users at $50,000 -- states that the platform subscription is charged per app integration, and states that both tiers additionally require a separate per-user-per-year content SKU priced by content type across video, voice and chat. That per-user rate is the part that scales with the organisation and it is not published anywhere located, which is what keeps this off A along with silence on implementation. The vendor's own website carries no pricing page at all: every call to action on thetalake.com is a demo request. The refund position is published on the same listing, with fees non-cancellable and non-refundable except as required by law. Also recorded: RingCentral MVP customers are stated on the vendor's site to receive advanced archiving and eDiscovery capability through that reseller relationship. AWS Marketplace listing read 12 September 2026.
Firm and Practice Coverage
Who the product is actually built for. AmLaw, midlaw, small firm and solo, in house departments, government and courts, and which practice areas are supported rather than merely claimed.
Segments and the regulatory regimes served are described with substance, short of limits on the AI, which is the B band. The vendor serves wealth management, broker-dealers, RIAs and banks, public sector bodies handling FOIA, energy and utilities under FERC, NERC and CFTC oversight, and life sciences, with separate small and mid-sized and enterprise offerings, and names SEC Rule 17a-4 and FINRA supervision among the rules it supports. The agreement sets use limits, such as capturing only employees' communications. Nothing states which languages, channels or misconduct types the AI agents handle less well. Verified 18 September 2026.
Who this serves is described with real substance and the boundaries are left open, which is B. Industries are named and each carries its own page: financial services, state and local government, healthcare and telemedicine, education, and manufacturing and technology. The buying functions are named directly in the product material rather than as a tricolon: security and governance teams, compliance teams, retention and technology teams, and legal and eDiscovery teams, the last with specific promises about preservation, evidence sets and defensible productions. Regulatory coverage is specific, with pages for SEC 17a-4, CFTC 1.31, MiFID II, GDPR, HIPAA and CCPA. R15 applies to two limbs: law firm segment sizing and practice-area support do not bite on a platform bought by an enterprise compliance and legal function rather than by a practice group, and are neither credited nor penalised. What is not stated is where coverage stops: no jurisdictional or platform limits are published and nothing addresses what the product does not support. Verified 12 September 2026.
The 12 legal signals, side by side
Recorded rather than graded. These are the questions a practitioner has to answer before a tool touches a client matter, and the answers are taken from public material only.
Client Data in Training
Can material a lawyer puts into this product be used to train a model?
The published Services Agreement grants a use right over client data bounded to support and improvement of the services, and never names training. Section 4.2 licenses Smarsh to access and use client data as necessary to provide support and improve the services on the client's behalf. No published term addresses whether client communications train or adapt the models behind Smarsh's AI agents.
The Master Services Agreement, updated 12 March 2026, confines Theta Lake's use of Customer Data to three enumerated purposes tied to service provision -- providing the Service, generating Output and Reports for the customer, and creating Usage Data and Anonymized Data -- under a limited license running only for the Subscription Period, and the DPA's United States schedule states the vendor will not retain, use, disclose, sell or share Personal Data other than to provide the Services on the customer's documented instructions.
No surface located names training in either direction: not the agreement, not the DPA of 2 July 2025, not the privacy policy updated 8 January 2026, not any product page. Two qualifiers belong on the record. Anonymized Data is defined as data derived from Customer Data with all personal identifiers removed and then aggregated, is expressly not Customer Data, and may be used to improve and develop the Service, with the customer able to opt out of its creation and use by emailing the vendor's legal address.
Usage Data is telemetry, anonymized and aggregated, and is also used to improve and develop the Service. Neither clause names machine learning, models or training, so neither is recorded as a training permission.
Prompt and Output Retention
How long does the product keep what a lawyer typed, and can that be set to zero?
The client sets retention. The Services Agreement retains archived data for client-set periods (up to seven years by default, longer for a fee), keeps capture data for a client-configured temporary period of up to 30 days before deletion, and provides for deletion after termination, though one section says as soon as practicable and another allows up to six months. Retention of AI summaries and prompts is not addressed separately.
Retention is customer-set and no retention is an available setting. The AI Communication and Interaction Governance page states that customers can selectively collect and apply dynamic archiving retention rules to decide which records to keep, for how long and in what region, as well as what to explicitly not retain, alongside WORM with 17a-4 attestation options. The privacy policy states customers control retention settings in the Services and can apply multiple retention periods to their data, and DPA Annex II records limited data retention on durations defined by the data exporter.
Deletion is separately committed in DPA section 8. The material retained is captured communications, AI interaction records and generated Reports rather than a lawyer's prompts to a drafting assistant, which is the shape this signal takes on a communications archive; the vendor also states it retains backups for business continuity and disaster recovery.
Ethical Walls and Matter Segregation
Does retrieval respect the firm’s ethical walls, or can the model read across them?
Role-based access is described without published detail on separating matters or investigations. The agreement lets clients set user roles with different access levels, and the innovations FAQ refers to role-based access and chain of custody around AI outputs. How access is walled between investigations or legal holds is not documented in the materials read.
The product runs its own permission model rather than enforcing a document management system's access model at query time. What is documented: role-based access control, including RBAC applied to private links into contextual investigation views, group and role-based policy notifications, SAML single sign-on federated to the customer's identity provider, two-factor authentication, and DPA Annex II measures for user identification and authorization.
The trust center carries separate access control, data access, access monitoring and logging entries, all behind an access request. No ethical wall, conflicts check or matter-level segregation construct was located on any surface: legal hold cases exist as a workflow object but are not described as an access boundary, so a firm would have to keep the product's roles aligned with its own walls itself.
Third Party Request and Subpoena Notice
If someone subpoenas the vendor for a firm’s data, does the firm hear about it first?
The Services Agreement commits to reasonable notice before compelled disclosure of confidential information, including client data, where feasible and legally permitted, so the client can contest the order, and to cooperate at the client's expense.
Both instruments commit to notice and neither publishes a transparency report. MSA section 8(c) permits disclosure of the other party's confidential information where required by law including by court subpoena, but only where written notice is given first so the disclosing party can contest the disclosure, seek to limit it or obtain a protective order, and requires that only the legally required portion be furnished with confidential treatment sought for it; Customer Data, Reports and Output are the customer's Confidential Information under section 1(d).
DPA section 6(b) commits that on receipt of a binding public authority order for Personal Data the vendor will notify the customer unless legally prohibited, and Schedule 1 section 4(l) directs SCC clause 15 notification to the customer rather than to data subjects. No transparency report and no figures on requests received were located on any surface, which is what separates this from the top value.
Primary Law Corpus Provenance
Where does the law in this product come from, and does the vendor have the right to use it?
Searched the innovations and trust pages and the Services Agreement on 18 September 2026. The AI works over the client's own captured communications; no external legal corpus is described.
No primary law corpus is identified because the product does not use one. This platform answers from the customer's own captured communications -- voice, video, chat, mobile messaging, email and whiteboard content ingested through certified API integrations -- rather than from case law or statute, so the provenance and licensing question does not arise in the form the signal asks it. What stands in its place is a published detection policy library of more than 80 built-in and custom policies mapped to named regimes including SEC 17a-4, CFTC 1.31, MiFID II, GDPR and HIPAA, each carrying its own regulation page. No third-party legal corpus is used, so no licensing basis is owed.
Good Law Verification
Does the product tell you when the authority it just cited has been overruled?
Searched the same surfaces on 18 September 2026. The product does not cite legal authority, so no subsequent-history check arises and none is described.
The product does not cite legal authority, so no located material addresses checking subsequent history, and none would be expected of it. Detections reference the customer's own captured content and the vendor's policy library rather than decided cases. Recorded so the row states the position rather than leaving a reader to infer it from silence.
Refusal and Uncertainty Behavior
What does the product do when the answer is not in the corpus?
Searched the innovations page and FAQ, the trust page, the Services Agreement and the 2026 press releases on 18 September 2026. No abstention path, confidence score or grounding indicator is described for AI summaries or risk signals; the agents suppress low-relevance alerts rather than flagging uncertain ones.
No located public material describes an abstention path or what the system does when it cannot ground an answer. The closest published material is MSA clause 10(d), which states that output is machine-generated, may contain errors and misstatements, may be incomplete or inaccurate, may differ from one use to the next, and is the customer's sole responsibility to evaluate including by human review: that allocates responsibility for a wrong answer rather than describing abstention behavior.
Detection output is surfaced as prioritized alerts into a review workflow, and no confidence or grounding score was located on any surface, so the confidence-signal value is not true of this record either.
Fabricated Citation Record
Does a public court record exist addressing fabricated or hallucinated legal citations in output from this product?
Searched the AI Hallucination Cases database maintained by Damien Charlotin and trade press reporting on 18 September 2026 for court records addressing fabricated or hallucinated content in output from Smarsh products. None located. This signal does not record litigation history of any other kind.
No court order, opinion or disciplinary record naming this product or Theta Lake, Inc. was located as of 12 September 2026. Searches were run on both the product name and the company name against published trackers of AI hallucination decisions, including coverage of the Charlotin AI Hallucination Cases database, and returned nothing involving this vendor. This is a statement about the public record on that date and not a finding about the product. The product does not generate legal citations, which is the conduct those records address.
Bar Guidance Alignment
Has the vendor engaged in public with the ethics opinions its buyers are bound by?
Searched the same surfaces on 18 September 2026. The vendor names securities regulations such as SEC Rule 17a-4 and FINRA supervision rules, which are regulatory recordkeeping and supervision requirements, but no bar ethics opinion or court rule on AI.
No located public material engages with bar or ethics guidance. The vendor publishes extensively on regulatory obligation -- SEC, CFTC, MiFID II, GDPR, HIPAA and state privacy law, each with its own page -- and an Ethics Policy is named on its trust center behind an access request, but a corporate ethics policy is not engagement with the professional responsibility guidance a lawyer buyer is bound by, and nothing located addresses ABA Formal Opinion 512 or any state bar opinion on generative AI. Recorded as of 12 September 2026.
Billing and Fee Posture
Does the vendor address what happens to the bill when the work takes an hour instead of six?
The product is bought by regulated firms for their own compliance and investigations, where no client is billed for the work. Savings are claimed for the buyer's own costs, including reduced outside counsel spend and investigation costs cut by up to 75 percent.
The product does not touch a fee between a lawyer and a client. It is bought by an enterprise compliance, security or in-house legal function to supervise and archive that organization's own communications, and no client is billed for the work the detection performs. Savings claims are published -- reduced alert fatigue, lower investigation and review cost, and the cost of over-retention -- and they are aimed at the buyer's own operating cost rather than at a client invoice, so under the value's own terms they are recorded here and do not make this a savings-claims row. R21 noted: this signal is specific to AI-assisted billable work, which this product does not produce.
Outside Counsel Guideline Readiness
Can a firm get this vendor through a client’s AI clause without a bespoke negotiation?
Searched the Services Agreement, trust page and legal documents index on 18 September 2026. The agreement lists subprocessors for mobile capture only (TeleMessage, Microsoft Azure, AWS and CallCabinet entities); no AI or model provider is named, and the DPA and Information Security Addendum are available on request rather than published.
A subprocessor list is published and contractually maintained, and the model provider limb is not met. DPA section 7 and Annex III place the list at a named URL on the vendor's support portal, require it to carry each subprocessor's name, location and the activities it performs, require the vendor to inform customers of intended additions or replacements, and give the customer an objection right with termination as the remedy if no resolution is reached.
The article itself was refused to this index's fetcher on 12 September 2026, so its contents were not read; the trust center carries a subprocessors entry behind an access request. Forwardable client-facing material does exist in the published DPA, which satisfies the third limb of the top value, but no statement of which model providers see customer content was located on any ungated surface -- a 'How We Use AI' document is named on the trust center behind the same request -- so the disclosure pack value is not available on the evidence.
Court Disclosure Support
If a judge’s standing order requires an AI disclosure, can the product produce one?
Some elements of a defensible record exist, short of an AI disclosure record. The vendor describes chain of custody, full audit trails, traceable source data and reproducible AI outputs, and the archive supports legal holds and exports. Nothing records which model produced a summary or risk signal, or who reviewed it.
Elements of a record exist, short of a document-level export built for a court's AI disclosure. The product captures AI interactions as records identifying which assistant was used along with the prompts and responses, reconciles them into an interaction timeline, applies legal hold to that content automatically for custodians in a matter, keeps review actions and inserted policy notifications in an audit history that evidences what users were told, generates conversation audit reports, and exports to the customer's own storage.
What it does not produce is a per-document certification tying a named model, the sources it retrieved and a named human verifier to a filing, because the record it holds is of the organization's communications rather than of a brief's drafting. A firm asked to evidence AI use in a matter would have material to draw on and would have to assemble the certification itself.
The questions both sides leave open
Derived from the records above rather than written, so it cannot favor either vendor. Take these into both conversations and ask each side the same question.
- Primary Law Corpus Provenance
- Good Law Verification
- Refusal and Uncertainty Behavior
- Bar Guidance Alignment
Which one fits
Choose Smarsh if
- Your auditors want the reports as part of the contract, not after a request. Smarsh's Services Agreement commits to annual independent audits under ISO 27001 or SSAE 18 and to give clients its most recent ISO 27001 and SSAE 18 reports and a penetration test summary as standard audit documentation.
- Your firm sits outside core financial services, or spans several regulated sectors. Smarsh serves wealth management, broker dealers, RIAs and banks, public bodies handling FOIA, energy and utilities under FERC, NERC and CFTC oversight, and life sciences, with capture from mobile carriers such as Verizon and AT&T and apps such as WhatsApp and Signal.
- Your legal team needs to narrow what goes to outside counsel. Smarsh's Discovery Agent produces summaries, timelines and custodian mapping over archived communications to cut what is exported, and the archive supports legal holds, with retention periods you set, up to seven years by default.
Choose Theta Lake if
- Your governance review wants the AI itself audited. Theta Lake holds ISO/IEC 42001:2023 certification for AI management alongside SOC 2 and PCI DSS, and its agreement states plainly that machine generated output may be inaccurate or differ between runs and must be evaluated by the customer, including by human review.
- You need to supervise how employees use AI assistants. Theta Lake captures interactions with Microsoft Copilot, Zoom AI Companion, Anthropic's Claude and OpenAI's tools as records, with classifiers for prompt injection, jailbreak behavior, shadow AI and sensitive information sharing, and applies legal hold to that content for custodians in a matter.
- You want control over what is kept, for how long and where. Theta Lake lets you set retention rules, choose not to retain some records at all, and store in the region you choose, with customer specific encryption keys you can manage in AWS or Azure and a documented Relativity integration for legal hold and export.
In summary
Smarsh
Smarsh, of Portland, Oregon, captures, archives, supervises and searches the communications of regulated firms, mainly in financial services, across email, messaging, collaboration tools, mobile including WhatsApp, voice and social media, to meet rules such as SEC Rule 17a-4 and FINRA supervision. Its AI agents suppress low relevance surveillance alerts, detect misconduct across languages, summarize and translate messages, and support investigations. The AI Legal Index grades it in the top two bands on ten of fifteen capability axes. Its published Services Agreement lets clients set retention, commits to notice before compelled disclosure and provides ISO 27001 and SSAE 18 reports to clients. As of 18 September 2026 the index located no statement on whether client data trains its models, no named model and no published price.
Theta Lake
Theta Lake, based in Santa Barbara, California, is a communications governance and archiving platform for regulated organizations, capturing voice, video, chat, mobile messaging, email and whiteboard content through more than a hundred API integrations and applying machine learning to detect compliance, conduct and data protection risk. It also governs employees' interactions with AI assistants. The AI Legal Index grades it in the top two bands on fourteen of fifteen capability axes, with an A on data stewardship. It holds ISO 42001 certification for AI management, and its published agreement and data processing addendum cover confidentiality, deletion, subprocessors and notice before compelled disclosure. As of 12 September 2026 the index located no measured detection accuracy, no named model and no published per user content price.
Questions buyers ask
Smarsh vs Theta Lake: which is better for communications compliance?
On published evidence Theta Lake sits in the top two bands on fourteen of fifteen AI Legal Index capability axes and Smarsh on ten of fifteen, and Theta Lake is level or higher on every axis. Its lead comes from how it documents its AI: ISO 42001 certification, detections linked to source content, and contract terms on deletion and subprocessors. Smarsh serves a broader range of regulated sectors and commits contractually to hand clients its audit reports.
Does Smarsh's AI suppress surveillance alerts without review?
Smarsh's Noise Reduction Agent and Intelligent Agent suppress low relevance alerts so supervisory teams see fewer, and Smarsh reports false positives down 60 percent without publishing a method. The vendor describes the agents as augmenting rather than replacing human expertise, with audit trails and chain of custody. Nothing published says whether suppressed alerts can be reviewed or sampled, what threshold governs suppression, or who approves the configuration. Graded by AI Legal Index against 15 capability axes and 12 legal signals, including privilege handling and citation accuracy, from each vendor's own published materials, verified September 25, 2026. No vendor pays for placement.
How much does Theta Lake cost?
Theta Lake's AWS Marketplace listing, where it is the seller of record, shows an SMB platform tier for up to 999 users at $15,000 a year and an Enterprise tier for 1,000 or more users at $50,000 a year, charged per app integration. Both tiers also need a per user per year content charge for video, voice or chat that is not published. Fees are non cancellable and non refundable. Smarsh licenses per connection and publishes no figure. Graded by AI Legal Index against 15 capability axes and 12 legal signals, including privilege handling and citation accuracy, from each vendor's own published materials, verified September 25, 2026. No vendor pays for placement.
Do Smarsh and Theta Lake train AI on client communications?
Neither says so in terms. Smarsh's Services Agreement lets it use client data to provide support and improve the services on the client's behalf, without naming training. Theta Lake's agreement confines use of customer data to providing the service, generating reports, and creating anonymized aggregated data that may improve the service, with an opt out by email, again without naming training. Both tie data use to the service, and neither states a position on model training. Graded by AI Legal Index against 15 capability axes and 12 legal signals, including privilege handling and citation accuracy, from each vendor's own published materials, verified September 25, 2026. No vendor pays for placement.
What do Smarsh and Theta Lake both leave unpublished?
The models and their error rates. Neither names the model behind its detection or says when it changes, and neither publishes a measured recall or precision figure for the risk it flags. Neither says what its AI does when it cannot judge a conversation with confidence. Neither engages with bar guidance for lawyers relying on the archive, and neither offers a record tying an AI finding to the model that produced it and the person who reviewed it. Graded by AI Legal Index against 15 capability axes and 12 legal signals, including privilege handling and citation accuracy, from each vendor's own published materials, verified September 25, 2026. No vendor pays for placement.
Two readings to weigh. Theta Lake's own pages disagree in places: its trust center lists CSA STAR Level 1 while its AI governance page claims Level 2 for AI, and it lists ISO 27001 as a compliance item while its security architecture page says only that controls are aligned with it. Smarsh publishes performance figures for its AI agents, such as 60 percent fewer false positives, without a sample or method, and its own terms conflict on whether client data is deleted as soon as practicable after termination or kept for up to six months. Smarsh was verified on 18 September 2026 and Theta Lake on 12 September 2026. Neither vendor reviewed this page.
Neither vendor paid for inclusion, placement or a grade, and neither reviewed this page before it published. Everything above comes from public material on the dates shown. How the index grades.