Every lawyer has heard some version of this story: a brief goes out, the citations look clean, and then opposing counsel or the judge discovers the cases don't exist. Mata v. Avianca made it infamous in 2023. It was never going to be a one-off, and the numbers back that up. Damien Charlotin's AI Hallucination Cases Database now tracks more than 1,600 court decisions involving AI hallucinations worldwide.
In a new piece for Legal Reader, Dan Ivtsan, Steno's Senior Director of AI Products, lays out why "our AI is accurate" isn't a claim any firm should take on faith, and how his team built a continuous benchmark to test it themselves.
Dan draws a distinction that's easy to miss:
Let’s be clear, a hallucinated citation isn’t a typo. It’s a fake authority presented with confidence in a profession where authority is everything. It comes in two flavors, one far sneakier than the other:
-
-
-
-
Fabrication - where the case doesn't exist at all. Embarrassing, but increasingly catchable: the citation simply doesn’t resolve.
-
Misattribution - the case is real, the citation is accurate, but the opinion doesn’t say what the model claims it says. This is the dangerous one, because it survives at a glance. The citation checks out; the proposition doesn’t.
The piece explains why public benchmarks and one-off human review both fall short, either because they go stale within months of publication or because lawyer-hour math makes large-scale manual verification impossible. Instead, Dan describes how his team built their own pipeline on top of Stanford RegLab's published legal research question set, verifying every citation against CourtListener's open case law database.
A citation either resolves to a real case or it doesn't. The opinion either supports the stated proposition or it doesn't. The case is either still good law or it's been overruled. These are checkable facts.
The result was roughly 1,800 individually verified citations across nine frontier models, sliced by question type, case age, model price point, and consistency across repeated runs. Dan's takeaway for firms evaluating legal AI is to stop asking vendors for accuracy claims and start asking for citation-level verification statistics, on what questions, verified against what source.
Read the full piece on Legal Reader.