The failure everyone talks about is the invented case. It is the least dangerous of the three, because it is the one you catch.
The other two are harder. A real citation attached to the wrong case name. And a real case, correctly named, cited for a proposition it does not support. Both arrive looking exactly like good work.
The risk is not hypothetical. The Courts and Tribunals Judiciary’s AI guidance warns that generated material may include fictitious cases, citations and quotes, and places responsibility for checking accuracy on the person using the output.
This is a routine for catching all three before anything leaves your desk. It takes about twenty minutes the first time and a few minutes thereafter, and it requires nothing beyond the research tools you already have.
Before you start: what you are actually checking
There are four possible states for any citation, and most people collapse them into two. Separating them is the whole technique.
| Verdict | Means |
|---|---|
| Verified | It resolves, and the case name matches the reporter entry |
| Name mismatch | It resolves, but to a different case than the one named |
| Unresolved | Not found in the database — which is not proof of fabrication |
| Not checked | Your lookup failed. This says nothing at all about the citation |
The fourth is the one that matters. If a database is slow, rate-limits, or you simply did not get to it, that is not a finding. Recording it as “unverified” quietly converts an absence of information into evidence, and that is how people end up dismissing citations that were perfectly good.
Keep a column for it. Literally a column — the discipline only works if the distinction is written down.
The routine
1. Extract every citation into a list. Do not verify in place, in the flow of the document. You will skim. Pull them into a separate list — a table with one row per citation — so each one gets looked at on its own.
2. Resolve the citation, not the name. Take the volume, reporter and page and look that up first, without the case name in front of you. Then compare what came back against what the document claims.
This ordering matters more than it sounds. If you search the case name, you will find the case, confirm it exists, and move on — never having checked whether the citation points at it. That is exactly the error a plausible fabrication is built to survive.
3. Mark the verdict, including “not checked.” One of the four above. No free text, no “looks fine.”
4. For anything that survives, check the proposition. This is the step no tool does for you. Open the case and confirm it says what your document says it says. Pinpoint citations are where this breaks down most often — the case is real, the page is real, and the passage on that page does not support the sentence it is attached to.
Budget most of your time here. Steps 1–3 are mechanical; this one is legal work.
The two examples worth internalising
A citation that resolves cleanly and is still wrong. Take a real reporter volume and page, attach a plausible-sounding party name to it, and you have something that passes any existence check. Volume real, reporter real, page real, resolves on the first try — and the case at that citation is a completely different matter. An existence-only check reports it verified and moves on. This is precisely the shape of what language models produce: plausible names attached to real numbers.
A real case with two digits transposed. Correctly named, correctly described, and the citation points at a different page or volume. A checker reports it identically to the citation that was invented outright. Nothing in the output distinguishes a typo from a fabrication.
Which is the honest limit of the whole exercise: verification produces a review queue, not a verdict. It tells you where to look. It does not tell you what you found.
Setting the threshold
If you use any fuzzy matching to compare a document’s case name against the reporter entry — and you should, because legitimate citations are shortened all the time — set the threshold low.
The reasoning is asymmetric. A missed mismatch costs you one more manual check. A false alarm on a legitimately shortened caption costs you the tool: flag Brown v. Board of Education against Brown v. Board of Ed. of Topeka twice and nobody will run it a third time. Tune for the failure you can live with.
What you should have at the end
A table, one row per citation, with a verdict against each and a separate column for the ones you could not check. That table is the artefact. It is what you keep, what you can hand to a supervising partner, and what you can point at in six months when somebody asks how the research was checked.
It is also the thing that makes the work defensible rather than merely done.
Sources and methodology
- Scope
- A cross-jurisdictional quality-control routine. Citation formats, authoritative databases and professional duties differ by court, regulator and jurisdiction; the applicable local process still governs.
- How this was produced
- Built from the author's legal research practice and systems-design work. The four-state ledger is a review protocol, not an automated determination that a legal authority is valid or supports a proposition.
- Artificial Intelligence Guidance for Judicial Office HoldersCourts and Tribunals Judiciary
- Risk Outlook report: The use of artificial intelligence in the legal marketSolicitors Regulation Authority
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and Technology
Read the editorial standards, corrections policy and AI-use disclosure.
