In June 2023 a federal judge in Manhattan fined two lawyers and their firm $5,000 for filing a brief built on opinions that did not exist. The opinions came from ChatGPT. The sanction in Mata v. Avianca was not for using the tool. It was for standing by the citations after the court asked about them, without ever having looked them up.
That is still the whole lesson. The ABA’s Formal Opinion 512 (July 2024) and the Texas ethics committee’s Opinion 705 (February 2025) say the same thing in different words: the lawyer is responsible for the work product regardless of who, or what, drafted it. Verification is not optional and it is not delegable to the tool that produced the draft.
Here is a routine that does it in about twenty minutes the first time and a few minutes thereafter. It catches the invented case, which is the easy one, and the two failures that arrive looking exactly like good work.
The three ways a citation goes wrong
- The case does not exist. Invented outright. The least dangerous of the three, because any existence check catches it.
- The citation is real but belongs to a different case. Volume, reporter and page all resolve. The case at that citation is not the one named. An existence-only check reports it as fine.
- The case is real, correctly named, and does not say that. The pinpoint points at a page that does not support the sentence attached to it, or the holding has since been overruled. No tool catches this. Only reading does.
Before you start: the four states
Most people collapse a cite-check into two outcomes, verified or not. There are four, and separating them is the whole technique.
| Verdict | Means |
|---|---|
| Verified | It resolves, and the case name matches the reporter entry |
| Name mismatch | It resolves, but to a different case than the one named |
| Unresolved | Not found in the database, which is not proof of fabrication |
| Not checked | Your lookup failed or you did not get to it. This says nothing about the citation |
The fourth is the one that matters. If a database is slow, rate-limits, or you ran out of time, that is not a finding. Recording it as “unverified” quietly converts an absence of information into evidence, and that is how good citations get struck from a brief. Keep a column for it. Literally a column. The discipline only works if the distinction is written down.
The routine
1. Pull every citation into a list. Do not verify in place, in the flow of the document. You will skim. One row per citation, in a table, so each one is looked at on its own.
2. Resolve the citation, not the name. Take the volume, reporter and page and look those up first, without the case name in front of you. Then compare what came back against what the brief claims.
The ordering matters more than it sounds. If you search the case name, you will find the case, confirm it exists, and move on, never having checked whether the citation points at it. That is exactly the error a plausible fabrication is built to survive.
Where to resolve: Westlaw or Lexis if the firm has them. If it does not, the CourtListener citation lookup from the Free Law Project is free. You can paste a whole brief (up to 64,000 characters) and it returns every citation it finds with a match status, about sixty citations a minute. Google Scholar’s case search works for one-at-a-time checks.
3. Mark the verdict, including “not checked”. One of the four above. No free text, no “looks fine”.
4. For anything that survives, check the proposition and the treatment. This is the step no tool does for you. Open the case and confirm it says what your brief says it says, at the page you cited. Then check the flag: KeyCite on Westlaw, Shepard’s on Lexis, or the citing decisions on CourtListener. A case can be real, correctly cited, accurately described, and overruled.
Pinpoint citations are where this breaks down most often. The case is real, the page is real, and the passage on that page does not support the sentence it is attached to. Budget most of your time here. Steps 1 to 3 are mechanical; this one is legal work.
5. Run it on anything a model touched, and ignore the model’s own assurances. A draft that says “all citations verified” has verified nothing. If a model produced or edited the research, every citation in it goes through steps 1 to 4 as if it had come from a stranger.
The two examples worth internalizing
A citation that resolves cleanly and is still wrong. Take a real reporter volume and page, attach a plausible-sounding party name, and you have something that passes any existence check. Volume real, reporter real, page real, resolves on the first try, and the case at that citation is a completely different matter. This is precisely the shape of what language models produce: plausible names attached to real numbers.
A real case with two digits transposed. Correctly named, correctly described, and the citation points at a different page or volume. A checker reports it identically to the citation that was invented outright. Nothing in the output distinguishes a typo from a fabrication.
Which is the honest limit of the whole exercise: verification produces a review queue, not a verdict. It tells you where to look. It does not tell you what you found.
Setting the threshold
If you use any fuzzy matching to compare the brief’s case name against the reporter entry, and you should, because legitimate citations are shortened all the time, set the threshold low.
The reasoning is asymmetric. A missed mismatch costs you one more manual check. A false alarm on a legitimately shortened caption costs you the tool: flag Brown v. Board of Education against Brown v. Board of Ed. of Topeka twice and nobody will run it a third time. Tune for the failure you can live with.
The checklist
Copy this into the matter file and fill it in for every brief.
- Every citation extracted to a table, one per row
- Each citation resolved by volume, reporter and page, before the name
- Verdict recorded from the four states; “not checked” kept separate
- Every surviving citation opened and the pinpoint read against the sentence
- Treatment checked (KeyCite, Shepard’s, or citing decisions)
- Anything a model drafted or edited treated as unverified until checked
- Table saved with the draft, with the checker’s name and the date
What you should have at the end
A table, one row per citation, with a verdict against each and a separate column for the ones you could not check. That table is the artifact. It is what you keep, what you hand to the supervising partner, and what you point at when somebody asks, months later, how the research was checked.
It is also the thing that makes the work defensible rather than merely done. The lawyers in Mata did not lack a tool. They lacked the table.
Sources and methodology
- Scope
- A quality-control routine for US practice, with the ethics authorities cited from the ABA and Texas. Citation formats, databases and professional duties differ by court and state; the local rules still govern. The routine itself works in any jurisdiction.
- How this was produced
- Built from the author's legal research practice and from building and evaluating a citation verifier against a labeled test set. The four-state ledger is a review protocol, not an automated determination that an authority is valid or supports a proposition.
- Mata v. Avianca, Inc., No. 1:22-cv-01461, Opinion and Order on Sanctions (S.D.N.Y. June 22, 2023)Justia
- Formal Opinion 512: Generative Artificial Intelligence ToolsAmerican Bar Association, Standing Committee on Ethics and Professional Responsibility
- Opinion 705: Lawyers' use of generative artificial intelligenceProfessional Ethics Committee for the State Bar of Texas
- Citation Lookup and Verification APIFree Law Project, CourtListener
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and Technology
Read the editorial standards, corrections policy and AI-use disclosure.
