The problem no one is measuring

Frontier models are confident, expensive — and quietly wrong.

They never say "I'm not sure." A fabricated citation gets a lawyer sanctioned. A wrong figure loses a deal. And you pay premium token rates for the privilege.

56%
of legal citations a cheap frontier model produced were mis-cited or fabricated whole cases
7–38%
bad-citation rate even for a top frontier model, depending on question difficulty
33–67%
of financial answers were silently wrong when source inputs were imperfect
9 of 10
fabricated citations that a “does it exist?” check waves through — because the number is real and only the attribution is wrong

Measured on our own benchmark suite against public registries (CourtListener, SEC EDGAR). See Evidence. Real-world reference: Mata v. Avianca — lawyers sanctioned for ChatGPT-fabricated case citations.

Advisers working through a confidential transaction timeline
Professional services — published reports carry the same exposure.

AttestAlly — because “the AI cited it” isn’t a defence.

This is not hypothetical

Courts are already sanctioning people for this.

A researcher at HEC Paris maintains a public database of court decisions involving AI-fabricated content. As of July 2026 it tracks 1,762 cases across 39 countries301 of them with monetary penalties. Every example below was read at the court's own docket or the journal's own retraction notice.

⚖️

Legal — sanctions, revoked admissions, disqualification

Mata v. Avianca (S.D.N.Y., Judge P. Kevin Castel, 2023): six fabricated opinions, $5,000 and a bad-faith finding. “A fake opinion is not ‘existing law’.”

Wadsworth v. Walmart (D. Wyo., 2025): nine cases cited, eight didn't exist — generated by the firm's own in-house AI. $3,000 and pro hac vice revoked.

Johnson v. Dunn (N.D. Ala., 2025): counsel disqualified and ordered to serve the sanction order on every client and every judge in every case they had. “Time is telling us — quickly and loudly — that those sanctions are insufficient deterrents.”

Coomer v. Lindell (D. Colo., 2025): ~30 defective citations; the “corrected” brief had the same errors.
THE COURT: “And did you double-check any of these citations…?” — “Your Honor, I personally did not check it.”

🧬

Medical — and the audit trail deleted

An AP investigation found a Whisper-based clinical transcription tool used by 30,000+ clinicians across 40 health systems (~7 million visits) — which erases the original audio, so transcripts can never be checked against what was actually said. It invented a drug that doesn't exist.

“You can't catch errors if you take away the ground truth.”
— William Saunders, former OpenAI engineer

A paper in Intensive Care Medicine was retracted after AI was asked to do something trivial — convert a list of PubMed IDs into a reference list. It invented sources instead. Only 5 of 15 references existed. One fabricated reference cited the journal itself.

🎓

Academic — it survives peer review

A case report published in a peer-reviewed radiology journal contained this sentence, which made it through peer review, copy-editing and typesetting:

“I'm very sorry, but I don't have access to real-time information or patient-specific data, as I am an AI language model.”

And when undisclosed AI content is found, it usually stays: of 768 documented cases in published literature, only 2.9% (23) were ever formally corrected or retracted.

📊

Finance & software — liability and attack surface

Moffatt v. Air Canada (2024): the airline argued it wasn't responsible for what its own chatbot told a grieving customer. The tribunal disagreed.
“It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot.”

Software supply chain (USENIX Security 2025): code models invented 205,474 unique package names — and 43% of them recur in all ten runs. “Not simply random errors, but a repeatable phenomenon.” A predictable hallucination is a registrable attack target — attackers now squat the names the models make up.

Sources: court dockets (RECAP/GovInfo/judiciary.uk), retraction notices, AP, and peer-reviewed papers — each verified at the primary document. We dropped nine widely-repeated “facts” about these cases that turned out to be wrong when we opened the source, including a misattributed judge's name and a quote Air Canada never said.

A drug-development team at work in a research laboratory
Research writing — fabricated references survive a careful read.
Checking a report against its own source pack

What we measured

And the weakness, since this page publishes those too: our limit is recall, not precision. We extract fewer figures from a long report than a careful human would. Everything we check, we check well; we do not yet claim to find everything there is to check.

Corroboration against your own documents — not verification against an external register.

The evidence

We tested this — hard — and published the failures too.

Multiple verticals, multiple registries, and a cross-lab panel of leading models. Where we claim a zero, we say exactly which set it was measured on — and where our own results don't flatter us, they're here too.

The Aggregate Index — 6 leading models, 5 labs

self-selected citations · same registry verifier for every model

Same task, same registry verifier, six leading models across five labs. We publish the results that don't favor us, too.

Medical & scientific
28% – 100% bad, every lab
0
Academic / scholarly
17% – 79% bad, every lab
0
Finance (figures)
67% – 87% of figures wrong*
0
Legal (case law)
3% – 58% (model-dependent)
0
Security (CVEs)
0–8%

*Finance — read this carefully: the wrong figures are overwhelmingly under-statements (89% of them, median 33% off). That's stale training data confidently presented as current — not invention. It is still a wrong number shipped with zero warning (exactly how a real report we audited stated a $1.5T company at $800B), but we won't call it hallucination, because it isn't.
Security — the result that doesn't favor us: frontier models handled well-known CVEs almost cleanly (0–8%), nothing like the other domains. One of the few flags was our own verifier being too strict on a terse registry entry. We're publishing that anyway.
One model refused. Asked for live market figures it couldn't ground, it declined rather than guess — the same thing our gate does. Not every frontier model fails; the ones that abstain are doing the right thing.
Our 0 is structural, not a winning streak: the gate emits only what a registry confirms. Named per-model results under NDA.

Don't take our word for it — independent, peer-reviewed research

every figure verified at the primary source

We didn't run these. Independent researchers did — and published them in peer-reviewed journals.

The tools by name — and the ones with no independent number at all

Tested with retrieval over their own proprietary case-law databases — the exact setup vendors say fixes hallucination.

ToolHallucination rateAccurate answers
Lexis+ AI17%65%
Westlaw AI-Assisted Research33%42% — accurate on fewer than half
GPT-4 (baseline)43%
HarveyNo independent number. Harvey reports ~1 in 500 — its own benchmark, its own scoring, its own model. Never third-party tested.
CoCounselNever independently tested. Thomson Reuters confirmed it was not the product Stanford evaluated.

Source: Magesh et al., J. Empirical Legal Studies (2025), 10.1111/jels.12413Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al. · J. Empirical Legal Studies · 2025Lexis+ AI hallucinated 17%, Westlaw AI 33% — WITH retrieval over their own case-law databases.Open the study →. When the vendor objected that Stanford tested the wrong product, Stanford re-ran on Westlaw AI-Assisted Research proper — it scored worse. An unaudited accuracy claim is a marketing number, not a control.

DomainWhat independent research measuredSource
Legal Leading AI legal-research tools hallucinate 17–33% of the time — with retrieval over their own proprietary case-law databases. The study notes a vendor had marketed “100% hallucination-free” citations. Magesh et al., J. Empirical Legal Studies (2025) · 10.1111/jels.12413Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al. · J. Empirical Legal Studies · 2025Lexis+ AI hallucinated 17%, Westlaw AI 33% — WITH retrieval over their own case-law databases.Open the study →
Journalism A newsroom (CNET) used AI to write 77 articles, then corrected 41 of them — a 53% correction rate; one explainer had at least five substantive errors. A separate outlet published articles under AI-generated fake author names. CNN & Washington Post (2023) · read at the primary outlet
Academic
measured on 2023-era models
55% of GPT-3.5 and 18% of GPT-4 citations were fabricated — and of the citations that were real, 43% / 24% still contained substantive errors. The number that ages well isn't the 55% — it's that the best model of its day still fabricated 18%. Walters & Wilder, Scientific Reports (2023) · 10.1038/s41598-023-41032-5Fabrication and errors in the bibliographic citations generated by ChatGPTWalters & Wilder · Scientific Reports · 2023GPT-3.5 fabricated 55% of citations, GPT-4 18%; of the real ones, 43% / 24% still had errors.Open the study →
Medical
measured on 2023–24-era models
Hallucinated references for systematic reviews: 28.6% (GPT-4), 39.6% (GPT-3.5), 91.4% (Bard). Chelli et al., JMIR (2024) · 10.2196/53164Hallucination rates and reference accuracy of LLMs for systematic reviewsChelli et al. · JMIR · 2024Hallucinated references: 28.6% (GPT-4), 39.6% (GPT-3.5), 91.4% (Bard).Open the study →
Security Code-generating models invented software packages that don't exist: ≥5.2% (commercial models), 21.7% (open-source) — 205,474 unique fabricated package names. Spracklen et al., USENIX Security (2025) · arXiv:2406.10279We Have a Package for You! — package hallucinations in code LLMsSpracklen et al. · USENIX Security · 2025Models invented 205,474 unique non-existent packages; ≥5.2% commercial, 21.7% open-source.Open the study →
AttestAlly, on 791 adversarial citation cases across legal, medical and academic: 0 false confirmations. That's the set it was measured on — not a claim about every document ever written. We certify only what resolves against a source of record, and flag the rest for you.

Honest scope: only the legal study tested retrieval-augmented professional tools — the others tested general-purpose models. The medical study covers a single clinical domain. These figures are date-stamped on purpose: they measure the models of their moment, and models improve — quoting a 2023 rate as if it were today's would make this page an example of the stale, unverified claim we sell against. We report these as the researchers did, and we don't stretch them. The finding that matters most to us is the academic one: a citation being real is not the same as it being right — which is exactly what an “does it exist?” check misses and our gate is built to catch.

We hold ourselves to the same standard

The number in that table was almost wrong. Here's how we caught it.

While building this page, our own research told the founder that the Stanford study found a 34% hallucination rate for one tool. It's a real study. It's a real finding. The number was wrong.

The paper says 33% — and states a range, "between 17% and 33%." The 34% came from a press release about the study, not the study itself. It had been copied across the internet, into blog posts, into summaries, into us. Nobody in that chain opened the paper.

We caught it exactly the way we catch our customers' errors: we went to the source of record before publishing. We pulled the full text, searched it, and found that "34%" appears nowhere in it. So the table above says 17–33%.

That's the whole product in one story. A real source. A plausible number. A confident chain of repetition. And a silent error that only a check against the original could catch — which is why we'd rather show you our own near-miss than pretend we don't have them. We're not 100%. We're honest about which 100% we're not.

Legal citation integrity — bad citations shipped

Frontier — cheap model
56.5%
56.5%
Frontier — top model (hard Qs)
38%
38%
"Does the case exist?" checker
48%
48%
AttestAlly gate
0

366 citations verified across 18 practice areas. An existence-only checker still ships 48% of mis-attributed cites (real citation, wrong case) — AttestAlly checks it's the right case: 0 of 198 slipped through.

Financial accuracy on imperfect source data — silent errors

Frontier (dirty inputs)
33–67%
≤67%
AttestAlly (dirty inputs)
0 / 6,000

Independent second-registry cross-check confirmed 354/354. Composed multi-step error bound: 0 bad of 5,048. Frontier failures concentrate on exactly the most complex, highest-value tasks.

The hard test: mis-attribution we build to beat ourselves

The dangerous error isn't a made-up citation — it's a real identifier attached to the wrong record. It looks authoritative and passes any “does it exist?” check. So we construct sets of exactly that, including same-topic collisions (two different records that share the same subject matter), and try to fool our own gate.

Legal (case law)
0 / 198
Medical (PubMed)
0 / 299
Academic (Crossref)
0 / 294
Pooled
0 false confirmations
0 / 791

What this does and doesn't say. An existence-only checker cannot detect any of these by design — every identifier in the set is real, so it confirms all of them. That's arithmetic, not a benchmark, and we won't dress it up as one. What the 0/791 says is narrower and more useful: when a real identifier is attached to the wrong record, our gate did not certify it.
Method note. These sets are adversarial and built against ourselves — the whole point is to lose. One early run did: it surfaced a real false-confirmation in our own matcher. We treated it as a design flaw rather than a tuning problem, fixed the underlying cause, made that exact case a permanent regression test, and re-ran clean — while confirming the fix hadn't simply bought a zero by refusing more, which would be a worse product wearing a better number. Our matcher errs toward abstaining, so where we report a frontier error rate, we are if anything understating our own strictness. Coverage cost is real: the remainder of each set is honest abstention, not silent success.

TestFrontier resultAttestAlly
Legal — cheap model, obscure citations56.5% bad (whole cases fabricated)0 shipped
Legal — top model, forced hard questions38% mis-cited0 shipped
Legal — adversarial mis-attribution (198 cases)48% (existence-only checker)0 / 198
Finance — multi-step analysis, dirty inputs33–67% silently wrong0 / 6,000
Finance — independent registry cross-check354 / 354
Calibration self-check (labeled set)0 false-confirm / 0 false-refute

All figures from AttestAlly's internal benchmark suite, verified against public registries. Methodology available under NDA.

Why our approach works — real examples

The silent error, caught.

Each of these was produced confidently, with zero warnings, by a frontier model. Each would have shipped.

✕ Frontier output — shipped confidently

"As established in Gunn v. Minton, 569 U.S. 251, federal jurisdiction requires…"

No warning. Looks authoritative. Cited in the brief.

✓ AttestAlly verdict

Mis-cited. 569 U.S. 251 resolves to Dan's City Used Cars v. Pelkey — a different case. The proposition's real support is elsewhere.

Corrected + flagged, with the CourtListener record cited.

✕ Frontier output — shipped confidently

"Total revenue for FY2023 was $8.24B, up 12% year over year…"

Plausible. Consistent-sounding. Feeds the valuation.

✓ AttestAlly verdict

Does not match the filing. The 10-K reports a different figure. The downstream 12% growth claim inherits the error.

Corrected to the SEC EDGAR value, with the source line cited.

✕ Frontier output — shipped confidently

"Studies show the treatment reduced mortality by 44% (Henderson et al., 2019)."

Confident statistic. Named source. Reads as settled.

✓ AttestAlly verdict

Unverifiable. No matching study/figure in PubMed. We do not certify it — we flag it for human review rather than guess.

Flagged, not silently kept. That's the discipline.

Examples are representative of documented failure modes in our benchmark data. Your document's findings are specific to it.

Real sanctions, real orders. Every figure on this page traces to a published court decision — the same standard we hold your document to.
See the cases behind these figures →

Three failures of the same assumption — that a careful reader is enough.

THE ATTESTALLY LEDGER S.D.N.Y. · 22 June 2023
COURT RECORD

Six cases.
None of them existed.

A $5,000 Rule 11 sanction after a brief cited six decisions that were never written. Asked to confirm the citations were real, the model said they were.

Read the original: CNBC, June 2023 →
THE ATTESTALLY LEDGER S.D.N.Y. Bankr. · 18 April 2026
THE FIRM DID NOT CATCH IT

Wall Street firm apologises to the court.

Sullivan & Cromwell filed wrong and fabricated citations in the Prince Group wind-down. Opposing counsel found them — not the firm’s own review.

Read the original: Bloomberg Law, April 2026 →
THE ATTESTALLY LEDGER Canberra · October 2025
NOT ONLY LAWYERS

A quote from a judgment that was never said.

A Deloitte report on a AUD 440,000 government contract carried a fabricated Federal Court quote and references to papers that do not exist. A university researcher found it.

Read the original: AP / Fast Company, Oct 2025 →

Every documented case we track →

Summaries by AttestAlly. Each clipping links to the original reporting.