They never say "I'm not sure." A fabricated citation gets a lawyer sanctioned. A wrong figure loses a deal. And you pay premium token rates for the privilege.
Measured on our own benchmark suite against public registries (CourtListener, SEC EDGAR). See Evidence. Real-world reference: Mata v. Avianca — lawyers sanctioned for ChatGPT-fabricated case citations.

AttestAlly — because “the AI cited it” isn’t a defence.
A researcher at HEC Paris maintains a public database of court decisions involving AI-fabricated content. As of July 2026 it tracks 1,762 cases across 39 countries — 301 of them with monetary penalties. Every example below was read at the court's own docket or the journal's own retraction notice.
Mata v. Avianca (S.D.N.Y., Judge P. Kevin Castel, 2023): six fabricated opinions, $5,000 and a bad-faith finding. “A fake opinion is not ‘existing law’.”
Wadsworth v. Walmart (D. Wyo., 2025): nine cases cited, eight didn't exist — generated by the firm's own in-house AI. $3,000 and pro hac vice revoked.
Johnson v. Dunn (N.D. Ala., 2025): counsel disqualified and ordered to serve the sanction order on every client and every judge in every case they had. “Time is telling us — quickly and loudly — that those sanctions are insufficient deterrents.”
Coomer v. Lindell (D. Colo., 2025): ~30 defective citations; the “corrected” brief had the same errors.
THE COURT: “And did you double-check any of these citations…?” — “Your Honor, I personally did not check it.”
An AP investigation found a Whisper-based clinical transcription tool used by 30,000+ clinicians across 40 health systems (~7 million visits) — which erases the original audio, so transcripts can never be checked against what was actually said. It invented a drug that doesn't exist.
“You can't catch errors if you take away the ground truth.”
— William Saunders, former OpenAI engineer
A paper in Intensive Care Medicine was retracted after AI was asked to do something trivial — convert a list of PubMed IDs into a reference list. It invented sources instead. Only 5 of 15 references existed. One fabricated reference cited the journal itself.
A case report published in a peer-reviewed radiology journal contained this sentence, which made it through peer review, copy-editing and typesetting:
“I'm very sorry, but I don't have access to real-time information or patient-specific data, as I am an AI language model.”
And when undisclosed AI content is found, it usually stays: of 768 documented cases in published literature, only 2.9% (23) were ever formally corrected or retracted.
Moffatt v. Air Canada (2024): the airline argued it wasn't responsible for what its own chatbot told a grieving customer. The tribunal disagreed.
“It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot.”
Software supply chain (USENIX Security 2025): code models invented 205,474 unique package names — and 43% of them recur in all ten runs. “Not simply random errors, but a repeatable phenomenon.” A predictable hallucination is a registrable attack target — attackers now squat the names the models make up.
Sources: court dockets (RECAP/GovInfo/judiciary.uk), retraction notices, AP, and peer-reviewed papers — each verified at the primary document. We dropped nine widely-repeated “facts” about these cases that turned out to be wrong when we opened the source, including a misattributed judge's name and a quote Air Canada never said.

And the weakness, since this page publishes those too: our limit is recall, not precision. We extract fewer figures from a long report than a careful human would. Everything we check, we check well; we do not yet claim to find everything there is to check.
Corroboration against your own documents — not verification against an external register.
Multiple verticals, multiple registries, and a cross-lab panel of leading models. Where we claim a zero, we say exactly which set it was measured on — and where our own results don't flatter us, they're here too.
Same task, same registry verifier, six leading models across five labs. We publish the results that don't favor us, too.
*Finance — read this carefully: the wrong figures are overwhelmingly under-statements (89% of them, median 33% off). That's stale training data confidently presented as current — not invention. It is still a wrong number shipped with zero warning (exactly how a real report we audited stated a $1.5T company at $800B), but we won't call it hallucination, because it isn't.
Security — the result that doesn't favor us: frontier models handled well-known CVEs almost cleanly (0–8%), nothing like the other domains. One of the few flags was our own verifier being too strict on a terse registry entry. We're publishing that anyway.
One model refused. Asked for live market figures it couldn't ground, it declined rather than guess — the same thing our gate does. Not every frontier model fails; the ones that abstain are doing the right thing.
Our 0 is structural, not a winning streak: the gate emits only what a registry confirms. Named per-model results under NDA.
We didn't run these. Independent researchers did — and published them in peer-reviewed journals.
Tested with retrieval over their own proprietary case-law databases — the exact setup vendors say fixes hallucination.
| Tool | Hallucination rate | Accurate answers |
|---|---|---|
| Lexis+ AI | 17% | 65% |
| Westlaw AI-Assisted Research | 33% | 42% — accurate on fewer than half |
| GPT-4 (baseline) | 43% | — |
| Harvey | No independent number. Harvey reports ~1 in 500 — its own benchmark, its own scoring, its own model. Never third-party tested. | |
| CoCounsel | Never independently tested. Thomson Reuters confirmed it was not the product Stanford evaluated. | |
Source: Magesh et al., J. Empirical Legal Studies (2025), 10.1111/jels.12413Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al. · J. Empirical Legal Studies · 2025Lexis+ AI hallucinated 17%, Westlaw AI 33% — WITH retrieval over their own case-law databases.Open the study →. When the vendor objected that Stanford tested the wrong product, Stanford re-ran on Westlaw AI-Assisted Research proper — it scored worse. An unaudited accuracy claim is a marketing number, not a control.
| Domain | What independent research measured | Source |
|---|---|---|
| Legal | Leading AI legal-research tools hallucinate 17–33% of the time — with retrieval over their own proprietary case-law databases. The study notes a vendor had marketed “100% hallucination-free” citations. | Magesh et al., J. Empirical Legal Studies (2025) · 10.1111/jels.12413Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al. · J. Empirical Legal Studies · 2025Lexis+ AI hallucinated 17%, Westlaw AI 33% — WITH retrieval over their own case-law databases.Open the study → |
| Journalism | A newsroom (CNET) used AI to write 77 articles, then corrected 41 of them — a 53% correction rate; one explainer had at least five substantive errors. A separate outlet published articles under AI-generated fake author names. | CNN & Washington Post (2023) · read at the primary outlet |
| Academic measured on 2023-era models |
55% of GPT-3.5 and 18% of GPT-4 citations were fabricated — and of the citations that were real, 43% / 24% still contained substantive errors. The number that ages well isn't the 55% — it's that the best model of its day still fabricated 18%. | Walters & Wilder, Scientific Reports (2023) · 10.1038/s41598-023-41032-5Fabrication and errors in the bibliographic citations generated by ChatGPTWalters & Wilder · Scientific Reports · 2023GPT-3.5 fabricated 55% of citations, GPT-4 18%; of the real ones, 43% / 24% still had errors.Open the study → |
| Medical measured on 2023–24-era models |
Hallucinated references for systematic reviews: 28.6% (GPT-4), 39.6% (GPT-3.5), 91.4% (Bard). | Chelli et al., JMIR (2024) · 10.2196/53164Hallucination rates and reference accuracy of LLMs for systematic reviewsChelli et al. · JMIR · 2024Hallucinated references: 28.6% (GPT-4), 39.6% (GPT-3.5), 91.4% (Bard).Open the study → |
| Security | Code-generating models invented software packages that don't exist: ≥5.2% (commercial models), 21.7% (open-source) — 205,474 unique fabricated package names. | Spracklen et al., USENIX Security (2025) · arXiv:2406.10279We Have a Package for You! — package hallucinations in code LLMsSpracklen et al. · USENIX Security · 2025Models invented 205,474 unique non-existent packages; ≥5.2% commercial, 21.7% open-source.Open the study → |
| AttestAlly, on 791 adversarial citation cases across legal, medical and academic: 0 false confirmations. That's the set it was measured on — not a claim about every document ever written. We certify only what resolves against a source of record, and flag the rest for you. | ||
Honest scope: only the legal study tested retrieval-augmented professional tools — the others tested general-purpose models. The medical study covers a single clinical domain. These figures are date-stamped on purpose: they measure the models of their moment, and models improve — quoting a 2023 rate as if it were today's would make this page an example of the stale, unverified claim we sell against. We report these as the researchers did, and we don't stretch them. The finding that matters most to us is the academic one: a citation being real is not the same as it being right — which is exactly what an “does it exist?” check misses and our gate is built to catch.
While building this page, our own research told the founder that the Stanford study found a 34% hallucination rate for one tool. It's a real study. It's a real finding. The number was wrong.
The paper says 33% — and states a range, "between 17% and 33%." The 34% came from a press release about the study, not the study itself. It had been copied across the internet, into blog posts, into summaries, into us. Nobody in that chain opened the paper.
We caught it exactly the way we catch our customers' errors: we went to the source of record before publishing. We pulled the full text, searched it, and found that "34%" appears nowhere in it. So the table above says 17–33%.
That's the whole product in one story. A real source. A plausible number. A confident chain of repetition. And a silent error that only a check against the original could catch — which is why we'd rather show you our own near-miss than pretend we don't have them. We're not 100%. We're honest about which 100% we're not.
366 citations verified across 18 practice areas. An existence-only checker still ships 48% of mis-attributed cites (real citation, wrong case) — AttestAlly checks it's the right case: 0 of 198 slipped through.
Independent second-registry cross-check confirmed 354/354. Composed multi-step error bound: 0 bad of 5,048. Frontier failures concentrate on exactly the most complex, highest-value tasks.
The dangerous error isn't a made-up citation — it's a real identifier attached to the wrong record. It looks authoritative and passes any “does it exist?” check. So we construct sets of exactly that, including same-topic collisions (two different records that share the same subject matter), and try to fool our own gate.
What this does and doesn't say. An existence-only checker cannot detect any of these by design — every identifier in the set is real, so it confirms all of them. That's arithmetic, not a benchmark, and we won't dress it up as one. What the 0/791 says is narrower and more useful: when a real identifier is attached to the wrong record, our gate did not certify it.
Method note. These sets are adversarial and built against ourselves — the whole point is to lose. One early run did: it surfaced a real false-confirmation in our own matcher. We treated it as a design flaw rather than a tuning problem, fixed the underlying cause, made that exact case a permanent regression test, and re-ran clean — while confirming the fix hadn't simply bought a zero by refusing more, which would be a worse product wearing a better number. Our matcher errs toward abstaining, so where we report a frontier error rate, we are if anything understating our own strictness. Coverage cost is real: the remainder of each set is honest abstention, not silent success.
| Test | Frontier result | AttestAlly |
|---|---|---|
| Legal — cheap model, obscure citations | 56.5% bad (whole cases fabricated) | 0 shipped |
| Legal — top model, forced hard questions | 38% mis-cited | 0 shipped |
| Legal — adversarial mis-attribution (198 cases) | 48% (existence-only checker) | 0 / 198 |
| Finance — multi-step analysis, dirty inputs | 33–67% silently wrong | 0 / 6,000 |
| Finance — independent registry cross-check | — | 354 / 354 |
| Calibration self-check (labeled set) | — | 0 false-confirm / 0 false-refute |
All figures from AttestAlly's internal benchmark suite, verified against public registries. Methodology available under NDA.
Each of these was produced confidently, with zero warnings, by a frontier model. Each would have shipped.
"As established in Gunn v. Minton, 569 U.S. 251, federal jurisdiction requires…"
No warning. Looks authoritative. Cited in the brief.
Mis-cited. 569 U.S. 251 resolves to Dan's City Used Cars v. Pelkey — a different case. The proposition's real support is elsewhere.
Corrected + flagged, with the CourtListener record cited.
"Total revenue for FY2023 was $8.24B, up 12% year over year…"
Plausible. Consistent-sounding. Feeds the valuation.
Does not match the filing. The 10-K reports a different figure. The downstream 12% growth claim inherits the error.
Corrected to the SEC EDGAR value, with the source line cited.
"Studies show the treatment reduced mortality by 44% (Henderson et al., 2019)."
Confident statistic. Named source. Reads as settled.
Unverifiable. No matching study/figure in PubMed. We do not certify it — we flag it for human review rather than guess.
Flagged, not silently kept. That's the discipline.
Examples are representative of documented failure modes in our benchmark data. Your document's findings are specific to it.
Three failures of the same assumption — that a careful reader is enough.
A $5,000 Rule 11 sanction after a brief cited six decisions that were never written. Asked to confirm the citations were real, the model said they were.
Sullivan & Cromwell filed wrong and fabricated citations in the Prince Group wind-down. Opposing counsel found them — not the firm’s own review.
A Deloitte report on a AUD 440,000 government contract carried a fabricated Federal Court quote and references to papers that do not exist. A university researcher found it.
Every documented case we track →
Summaries by AttestAlly. Each clipping links to the original reporting.