Legal · 2026-07-16 · 4 min read
The tools built to stop hallucination still hallucinate — with the numbers
Every legal-AI vendor says retrieval fixed the fabrication problem. The one peer-reviewed, preregistered study that tested them — with retrieval, over their own case-law databases — found otherwise.
17%Lexis+ AI hallucination rate
33%Westlaw AI hallucination rate
42%Westlaw AI accurate answers — fewer than half
When Thomson Reuters objected that Stanford had tested the wrong product, the researchers re-ran on Westlaw AI-Assisted Research proper. It scored worse. And the tools lawyers may trust most — Harvey, CoCounsel — have no independent numbers at all: Harvey's ~1-in-500 figure is its own benchmark, scored by Harvey, on Harvey's model. An unaudited accuracy claim is a marketing number, not a control.
The deeper problem isn't the made-up case. It's the real citation attached to the wrong case — it passes any "does it exist?" check. In our own testing, an existence-only checker waved through 9 of 10 such errors.
Source: Magesh et al., J. Empirical Legal Studies (2025)Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al. · JELS · 2025Lexis+ AI 17%, Westlaw AI 33% — with retrieval.Open the study →
Journalism · 2026-07-16 · 3 min read
A newsroom corrected 41 of 77 AI-written stories. That's the real risk.
CNET quietly used an AI tool to write 77 finance explainers. It then had to correct 41 of them — a 53% correction rate — with at least five substantive errors in a single compounding-interest piece.
For a science, health, or environmental desk, a fabricated study isn't a bad afternoon — it's a public correction, a credibility hit, and a story about you instead of your subject. The failure mode is specific and checkable: "a 2024 Nature study found…" where the DOI resolves, the journal is real, and the finding is misattributed. It reads as authoritative. It fails only against the actual paper.
That last category is exactly what a registry check catches — and exactly what a spell-check or a second AI pass does not.
Sources: CNN & Washington Post (2023), read at the primary outlet.
Method · 2026-07-16 · 2 min read
Why "just add more agents" doesn't fix a factual error
The intuitive fix for AI mistakes is more AI — a swarm of agents drafting, reviewing, and checking each other. For facts, it doesn't work.
When agents share the same underlying model, they share its blind spots. If the model has the wrong number, ten agents reviewing it reach the same wrong number — full confidence, still wrong. Consensus measures agreement, not truth. We watched a multi-agent pipeline produce a 40-page investment report that still shipped a market cap wrong by roughly 2×, with every agent confidently agreeing.
The only thing that breaks the tie is a source outside the models: the register of record. That's the whole idea.