Skip to content
Insights

Pilots die quietly

Nobody cancels a failed AI pilot. The licences renew, the champion moves practice group, and the usage report stops being circulated.

A failed AI pilot at a law firm does not produce a decision. It produces a silence.

Nobody calls a meeting to cancel it. The licences renew on the anniversary because renewal is the default and the invoice is small relative to everything else on the page. The partner who championed it moves practice group, or takes on a large matter, and stops asking about it. The usage dashboard that went round monthly goes round quarterly, then stops. Eighteen months later somebody asks whether the firm ever looked at that tool, and three people give three different answers.

This is the normal outcome. It is worth being precise about why, because the usual explanation — that the technology was not ready — is almost never the one that holds up.

This article does not claim a market-wide failure rate. It applies an operating view to the same lifecycle problem formalised by the NIST AI Risk Management Framework: governance, measurement and management have to continue after a system first works.

The pilot did not change how the work is produced

Most firm pilots are structured as access. A cohort gets licences, a training session, and a channel to ask questions in. Usage is measured. The theory is that capable people given a capable tool will find the value themselves.

For a genuinely general tool that theory is not unreasonable. It is how firms absorbed research databases and document comparison. But it has a specific failure mode, and the failure mode is this: nothing downstream of the lawyer’s own keyboard changed.

The precedent bank is unchanged. The matter opening process is unchanged. The review step before something leaves the firm is unchanged. The billing narrative is unchanged. The lawyer has a faster way to produce a first draft and no changed expectation about anything else. So the gain is absorbed entirely as personal convenience, which is real but invisible, and invisible gains do not survive a champion moving practice group.

The tools that survive in firms are the ones that changed a step. Document comparison survived because it replaced a task that had a name, an owner, and a place in a sequence. It was not adopted because people found it useful in general. It was adopted because “check the changes against the last version” became a thing the software did.

The billable hour is not the obstacle people think it is

The standard argument is that firms cannot adopt AI because AI reduces hours and hours are the product. This is half right in a way that makes it more misleading than being wrong outright.

The hour is the unit of price. It is not the unit of work. What actually constrains a firm is capacity — the number of matters it can carry at an acceptable standard with the people it has. A tool that compresses the input on a task does not automatically reduce revenue. It changes which of two things happens next: either the same work is done with less capacity consumed, or the same capacity now carries more work.

Firms that treat compression as a revenue threat get the first outcome and resent it. Firms that treat it as a capacity release get the second. The difference is not the technology. It is whether anybody decided in advance what the released capacity was for.

The harder economic question is not the hour. It is leverage. The associate pyramid works because a large volume of structured, checkable, learnable work sits underneath a small number of people who exercise judgement. A meaningful part of that work is exactly the kind that compresses first. A firm that compresses the base of its pyramid without deciding what juniors now do instead has not saved money. It has removed the mechanism by which its own seniors were produced, and it will not notice for about six years.

That is a strategy problem, not a procurement problem, and it is the reason this belongs at partnership level rather than with IT.

What a pilot has to answer

A pilot that only measures usage cannot tell you anything you can act on. High usage of a general tool tells you people tried it. Low usage tells you they did not. Neither says whether the work changed.

A pilot worth running is scoped to one workflow and answers four questions in order.

Which task, exactly. Not “research” or “drafting” — those are categories, not tasks. A task has an input, a decision, an output, a person who signs off, and a consequence if it is wrong. If you cannot write those five things down, the pilot has no object.

Who checks it, and against what. The check has to exist before the tool does, and it has to be somebody’s named responsibility rather than a general expectation of care. A review step that everybody owns is a review step nobody performs.

What the output is scored against. This is where most pilots quietly decline to participate. Scoring means a set of examples with known correct answers, produced independently of the system, against which the output can be marked. It is tedious to build and it is the only thing that converts an impression into a number. Without it you have a demonstration.

What happens to the capacity. Decided in advance, in writing. Otherwise the answer defaults to nothing, and nothing is what the pilot will have achieved.

The question that predicts survival

There is a single question that separates pilots that persist from pilots that fade, and it can be asked before anything is procured:

If this works, what stops being done the way it is done now?

If the answer is specific — this step, by this person, in this sequence — the pilot has somewhere to land. If the answer is that people will work faster, the pilot is a licence purchase with a training session attached, and its eventual quiet death is already scheduled. It just has not been circulated yet.

Sources and methodology

Scope
An operating perspective on law-firm AI adoption. It describes recurring delivery patterns and incentives; it does not claim a measured failure rate across the legal market.
How this was produced
Drawn from the author's legal, training and technical-project delivery experience, then tested against published risk-management guidance. No quantitative market dataset underlies the narrative examples.
  1. Risk Outlook report: The use of artificial intelligence in the legal marketSolicitors Regulation Authority
  2. AI Risk Management Framework CoreNational Institute of Standards and Technology
  3. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and Technology

Read the editorial standards, corrections policy and AI-use disclosure.