Do legal AI vendors train their models on client data?
There is no single answer, and that is the finding. Of the 108 legal AI vendors recorded on the Client Data in Training signal by the AI Legal Index, 8 prohibit training on customer content in the published agreement itself, 28 state it in a policy or trust page without a matching contract term, 3 train only where a customer affirmatively enables it, 1 train unless the customer turns it off, and 8 reserve the right to train in a published policy or agreement. The largest group, 60 vendors, publish no located term or policy addressing the question either way. Silence is a statement about disclosure rather than about the product, and it is also the one posture a firm cannot hold a vendor to.
This is the first question a general counsel asks and the one most often answered in a sales call rather than in a contract. The AI Legal Index records it on every vendor as the Client Data in Training signal, from published terms and policies, with a source and a verification date on each record. What decides the value is not whether the vendor says it respects confidentiality, it is where the commitment sits and whether a client can hold the firm to it.
Never, in the contract
Never, in the contractThe published agreement itself prohibits training on customer content. Not a policy page, the terms. This is the posture a firm can enforce, and it is the honest answer to which tools keep client data out of model training by default.
Never, in policy only
Never, in policy onlyA public policy or trust page states no training on customer content, with no matching term located in the published agreement. A real commitment, revisable unilaterally.
- AlexiLegal Research
- Alt LegalIP & Patents
- ClearbriefLitigation & eDiscovery
- CoCounsel LegalGeneral Legal Assistants
- DISCOLitigation & eDiscovery
- DraftwiseContract Review & Drafting
- EvePlaintiff & Claims AI
- EverlawLitigation & eDiscovery
- ExterroLitigation & eDiscovery
- GC AIGeneral Legal Assistants
- HarveyGeneral Legal Assistants
- LeahContract Review & Drafting
- LegalOnContract Review & Drafting
- LegalyzePlaintiff & Claims AI
- LegoraGeneral Legal Assistants
- Lex MachinaLegal Research
- LinkSquaresContract Review & Drafting
- LitifyPlaintiff & Claims AI
- MidpageLegal Research
- NextpointLitigation & eDiscovery
- NoxtuaGeneral Legal Assistants
- OntraLegal Ops & Spend
- PatlyticsIP & Patents
- QuestelIP & Patents
- Solve IntelligenceIP & Patents
- SpellbookContract Review & Drafting
- Streamline AILegal Ops & Spend
- WordsmithGeneral Legal Assistants
Opt in
Opt inTraining occurs only where the customer has affirmatively enabled it, which also leaves client data out of training by default.
Jhana.ai record an opt out posture, meaning training occurs unless the customer turns it off. On 7 vendors the recorded value is Permitted, in the contract, meaning the published agreement expressly reserves a right to train on customer content with no opt out located: Clio, Genie AI, Neos, Securiti, Sirion, SmartAdvocate and Unity ELM. LegalVIEW BillAnalyzer records Permitted, in policy only. Any de identification or aggregation qualifier a vendor attaches is recorded in the summary on its profile, because a qualified reservation is still a reservation.
And then the largest group: 60 of 108 vendors publish no located term or policy addressing the question either way. A signal records what public sources say on the date shown. It is not a grade and it is not a recommendation. Where a signal reads Not addressed, it means the index did not locate the material in public sources on that date, which is a statement about disclosure rather than about the product. For a buyer the practical consequence of silence is simple. There is nothing to enforce, so the answer arrives in a sales call and leaves no artifact behind.
Why the contract versus policy distinction decides this
A contractual prohibition sits in the terms a client can hold the firm to, so breaching it has a remedy. A policy statement is a public commitment the vendor can revise unilaterally. A marketing page that says your data is secure while the terms reserve a licence is a third thing again, and the gap between the three is where this risk actually lives. That is why the index records seven distinct values on this signal rather than a yes or a no, and why the tiers above are ordered by enforceability rather than by the warmth of the language.
The adjacent question, how long the product keeps what a lawyer typed, is recorded as Prompt and Output Retention on every vendor, and the full definition of every value on the training signal is on the signals reference.
Every record behind this page carries a source basis and a verification date, and a vendor that publishes a term tomorrow is redated the day it does. Standards and limits are on the methodology page. The accuracy side of the same diligence question is answered on which AI legal research tools publish their hallucination or accuracy rates, and the complete vendor set is in the directory.