Start a conversation Contact
← Research

Write down what would make you stop

Almost every AI programme can describe what success looks like. Very few can state the result that would end the work, and a project that cannot be stopped by evidence is not being evaluated — it is being funded.

There is a moment in most AI programmes, usually somewhere in the second year, when a reasonable person asks whether the thing is working. The answers come back and none of them is an answer. Adoption is up. The team has shipped. Users say it is helpful. A number is quoted from a pilot that ran on a different population. The question is not refused — it dissolves, because there was never a statement of what would have counted as it not working.

That absence is not an oversight. It is structural, and it is the same structure that makes a risk assessment worthless when its only possible outcome is approval. A test that cannot fail is not a test.

The claim: a stopping condition is the cheapest governance control available, and it has to be written before the result is known, because afterwards nobody can agree on one.

What a stopping condition is, and what it is not

It is not a risk register entry. It is not a phrase about monitoring. It is a sentence of the form: if this measurement, on this population, is worse than this value, we stop — and a named person who is accountable for acting on it.

Three properties make it real.

It is measurable with something that already exists, or that the project builds first. A condition that depends on data nobody collects is a condition that will never trigger.

It is checked on a schedule that is set in advance. Conditions checked when somebody remembers are checked when things are going well.

It names a consequence that is not “review”. Stop, roll back, restrict to a narrower population, or return to the previous process. “We will review it” is what a programme says instead of stopping.

None of this is novel as a principle. The Manage function of the NIST AI Risk Management Framework asks organisations to decide, on evidence, whether a system continues in service or is decommissioned. What is missing in practice is not the principle. It is the sentence that would let the decision be made by anyone other than the team whose work is being decided about.

The three stopping conditions every consequential system should carry Fig. 01

Under what measured result does this system stop, narrow or roll back?

  • The quality condition Accuracy on the held-back set falls below the threshold the process was designed around. Requires an evaluation set that exists before launch and is not used for tuning.
  • The harm condition Errors fall unevenly across groups, or a single failure exceeds a stated severity. Needs the outcome data disaggregated. If it is not collected, this condition cannot fire.
  • The value condition The measured benefit does not appear by the date it was forecast for. The one that is always omitted, and the only one that closes a project.

Why the third one is always missing

Quality and harm conditions get written, because governance frameworks ask for them. The value condition is the one nobody wants in the document, for a reason that is entirely human: it is the sentence that could end the programme that employs the person writing it.

So it has to be set by somebody else, and it has to be set at the point of funding rather than at the point of review. That is a small procedural change with a large effect — the same sponsor who approves the budget states the result that would mean the budget was wrong, and the statement is recorded with the approval rather than negotiated afterwards against a team’s reputation.

An organisation that does this discovers something uncomfortable and useful within about a year: most of its AI portfolio has never been evaluated against anything, and a portion of it can be stopped, which frees the capacity that the remaining portion actually needed.

What it does for the systems that survive

The argument for stopping conditions is usually made as a risk argument. The stronger argument is a commercial one.

A programme with a written stopping condition, checked on a schedule, has by construction a measurement, a baseline, a threshold and an owner. That is precisely the evidence a board asks for when it wants to expand something, and precisely what most successful AI projects cannot produce when asked to justify the next round of investment. The discipline that would have killed the project early is the discipline that makes the case for scaling it later.

What this does not tell you

A stopping condition does not make a system safe, and it is not a substitute for an impact assessment where one is required. It is the mechanism that gives an assessment somewhere to land — a way for a finding to have a consequence other than a paragraph.

It also does not mean that a triggered condition ends a programme automatically. It means the decision is made deliberately, by somebody accountable, with the evidence in front of them, rather than avoided by nobody looking. Overriding a stopping condition is a legitimate act. Overriding it silently is not, and the difference is whether the override is recorded.

The person who should act on this is whoever signs the next funding decision. Add one paragraph. Ask for the result that would mean this was the wrong thing to build, and the name of the person who has agreed to say so — because an assessment that can only conclude “proceed” is not an assessment, and a programme that can only conclude “continue” is not being managed.

Filed under · Method · Method · Governance · Decisions Inference Institute · 09 Jul 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.