Which AI legal research tools publish their hallucination or accuracy rates?
Of the 108 legal AI vendors graded by the AI Legal Index, 9 publish measured accuracy or hallucination evidence an outsider can check, earning an A on Citation Accuracy and Hallucination Disclosure: Alexi, CoCounsel Legal, Counsel Stack, Harvey, LegalOn, Midpage, PatSnap, Pre/Dicta and Vincent AI. A further 44 vendors earn a B, meaning grounding to primary sources and cite checking are documented without a published accuracy number. The remaining 55 either claim accuracy without documentation or publish nothing on the question at all.
Every vendor on this index carries a grade on Citation Accuracy and Hallucination Disclosure, built from public sources with a verification date. The grade measures the disclosure, not the marketing: an A requires measured evidence an outsider can check without contacting the vendor, a B requires documented grounding and cite checking short of a published number, a C records a claim made without documentation, and a D records that nothing was located. Legal has a documented public record of fabricated citations reaching filed briefs, so an untested claim of accuracy is not treated as evidence here.
Vendors with published accuracy evidence
An A on this axis means the accuracy or hallucination evidence is published, measured and checkable. Hover any grade for the exact band definition.
Documented grounding, no published number
A B on this axis means grounding to primary sources and cite checking are documented, and the measured accuracy number that would settle the question is not published.
- Bloomberg LawLegal Research
- BrightflagLegal Ops & Spend
- CasepointLitigation & eDiscovery
- ClearbriefLitigation & eDiscovery
- ClioIntake & Client Development
- CorlyticsRegulatory & Compliance Counsel
- CUBERegulatory & Compliance Counsel
- DefinelyContract Review & Drafting
- DescrybeLegal Research
- DigitalOwlPlaintiff & Claims AI
- DISCOLitigation & eDiscovery
- Docket AlarmLegal Research
- EvePlaintiff & Claims AI
- EvenUpPlaintiff & Claims AI
- EverlawLitigation & eDiscovery
- FilevinePlaintiff & Claims AI
- GC AIGeneral Legal Assistants
- IP AuthorIP & Patents
- IPRallyIP & Patents
- Jhana.aiLegal Research
- JosefLegal Ops & Spend
- LegalVIEW BillAnalyzerLegal Ops & Spend
- LegalyzePlaintiff & Claims AI
- LuminanceContract Review & Drafting
- MyCaseIntake & Client Development
- NeosPlaintiff & Claims AI
- Norm AiRegulatory & Compliance Counsel
- NoxtuaGeneral Legal Assistants
- PatlyticsIP & Patents
- Paxton AIGeneral Legal Assistants
- RegologyRegulatory & Compliance Counsel
- RelativityLitigation & eDiscovery
- RevealLitigation & eDiscovery
- SimpleLegalLegal Ops & Spend
- Solve IntelligenceIP & Patents
- SpellbookContract Review & Drafting
- StenoLitigation & eDiscovery
- SupioPlaintiff & Claims AI
- TavrnPlaintiff & Claims AI
- UniCourtLegal Research
- VixioRegulatory & Compliance Counsel
- WordsmithGeneral Legal Assistants
- Workday Contract Lifecycle ManagementContract Review & Drafting
- XLSCOUTIP & Patents
Read the shape, not just the names. The A tier is small because a published accuracy rate is a benchmark a competitor can beat and a plaintiff can quote, and most vendors have decided the safer move is a qualitative claim. The consequence for a buyer is that the loudest accuracy language in this market frequently attaches to the least measurable products, which is the exact inversion an index exists to catch.
Two signals sit alongside the grade on every vendor and answer the practical version of the question. Good Law Verification records whether a cited authority is checked for subsequent history, because a citation can be real, correctly quoted, on point, and dead. Refusal and Uncertainty Behaviour records what the product does when the answer is not in the corpus, because every fabricated citation that reached a filed brief passed through a moment when the correct output was that no support could be found. On the current corpus, 13 of 108 vendors surface treatment signals and 10 of 108 document an abstention or refusal path.
Do AI legal research tools cite real cases or hallucinate citations?
It depends on the product, and the honest version of the answer is checkable rather than reassuring. The AI Legal Index records this on every vendor. 13 of 108 vendors surface subsequent history treatment, from a licensed citator or a documented treatment signal of their own, so a cited authority is flagged when it is no longer good law. 10 of 108 document what the product does when it cannot ground an answer, which is the design decision every fabricated citation passed through. The index also records the Fabricated Citation Record signal, which tracks whether a public court record exists involving output from the product, with a citation required for any adverse value.
The court record on this question is public and growing, which is why this index bounds it carefully. The Fabricated Citation Record signal records what a court order actually says, distinguishes a filer’s verification failure from findings about a product’s own output, requires a docket or reporter citation for any adverse value, and treats nothing located as a statement about the record rather than a clearance.
An AI Legal Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. Grading standards and limits are published on the methodology page, and every vendor named here links to its full record across all fifteen axes and twelve signals. The confidentiality side of the same diligence question is answered on which legal AI vendors keep client data out of model training, and the complete vendor set is in the directory.