Request a scoping call Contact
← Research

Write down what would make you stop

Agree the evidence that would pause, narrow or end an AI use before approving it. Name the measurement, threshold and accountable owner, then connect review findings to a deliberate operating decision.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

An AI review needs evidence against the purpose for which the work was funded. Adoption, released features and favourable feedback can be useful observations without establishing the promised outcome. Agree the conditions for reconsideration before relying on those observations to justify continued spending.

If the acceptance process cannot produce a pause, restriction or rejection, its decision scope is incomplete. State the findings that would lead to each response. That makes an unfavourable result usable instead of leaving the review team to negotiate its meaning afterwards.

Set measurable reconsideration conditions before the result is known. They should identify the population, threshold, review time and person authorised to act. Conditions may later change, but the reason and evidence for that change must be recorded.

What a stopping condition is, and what it is not

It is not a risk register entry. It is not a phrase about monitoring. It is a sentence of the form: if this measurement, on this population, is worse than this value, we stop — and a named person who is accountable for acting on it.

Three properties make it real.

It is measurable with something that already exists, or that the project builds first. A condition that depends on data nobody collects is a condition that will never trigger.

It is checked on a schedule that is set in advance. Conditions checked when somebody remembers are checked when things are going well.

Name the immediate response to a breach: pause, restrict, roll back, escalate or enter a defined review state. A review can be appropriate if its authority, deadline and interim controls are explicit. An unspecified promise to review does not establish how the service behaves while the issue is unresolved.

None of this is novel as a principle. The Manage function of the NIST AI Risk Management Framework asks organisations to decide, on evidence, whether a system continues in service or is decommissioned. What is missing in practice is not the principle. It is the sentence that would let the decision be made by anyone other than the team whose work is being decided about.

The three stopping conditions every consequential system should carry Fig. 01

Under what measured result does this system stop, narrow or roll back?

  • The quality condition Accuracy on the held-back set falls below the threshold the process was designed around. Requires an evaluation set that exists before launch and is not used for tuning.
  • The harm condition Errors fall unevenly across groups, or a single failure exceeds a stated severity. Needs the outcome data disaggregated. If it is not collected, this condition cannot fire.
  • The value condition The measured benefit does not appear by the date it was forecast for. Assess the measured benefit against its agreed value and time conditions.
illustrative fixture

Choose the scope the evidence supports

Swipe or scroll for the full diagram →

Three deployment choicesAn accountable owner compares evidence with a named threshold, then expands, retains a narrower scope, or gathers more evidence.Evidence → threshold → ownerExpandNarrow scopeObserve
Three deployment choicesAn owner compares evidence with the threshold, then chooses expand, narrow scope or further observation.Evidence / thresholdExpandNarrow scopeObserve

What decision does the measured result support?

Owner: Accountable operating owner. Threshold: Record the quality, harm and value thresholds, target population and review date before collecting the result.

Expand
Expand only when the comparison clears every agreed threshold within its evaluated population.
Retain narrower scope
Retain the supported population or return to the incumbent when a guardrail fails.
Gather more evidence
Gather more evidence when the sample or outcome window cannot establish a decision.

Representative decision structure; the operating owner sets local thresholds, not this diagram.

Evidence and named thresholds lead to expansion, narrower scope or further observation.

Reviewed 2026-10-02

Include the value condition

Quality, harm and value require separate decisions. A service may perform accurately while failing to deliver an affordable benefit, or deliver a benefit while breaching a guardrail. Make those conditions visible to the accountable sponsor rather than compressing them into one success score.

The sponsor and relevant operating owners should agree the conditions at funding and acceptance. Include the evaluation team’s judgement about what can be measured. Preserve the approval record so later review can compare the outcome with the original assumptions.

Applying these conditions across a portfolio can identify work that should continue, narrow or stop. The result depends on the evidence for each use. It does not establish that most AI programmes are unevaluated or that a particular share should be cancelled.

What it does for the systems that survive

The argument for stopping conditions is usually made as a risk argument. The stronger argument is a commercial one.

A useful decision record connects the baseline, measurements, thresholds and owner. The same evidence can support expansion when outcomes meet the criteria. Merely writing a stopping condition does not establish that the data exists or that the service meets it, so test the review process as well.

What this does not tell you

A stopping condition does not make a system safe, and it is not a substitute for an impact assessment where one is required. It is the mechanism that gives an assessment somewhere to land — a way for a finding to have a consequence other than a paragraph.

A breach requires the agreed response and an accountable decision. Some conditions demand immediate suspension. Others permit a constrained review. Any exception needs authority, a documented reason, an expiry and interim controls. Recording an override does not by itself make it safe or lawful.

The funding owner should require the conditions for continuation, restriction and withdrawal with the initial proposal. Confirm that each can be measured and acted on. At review, accept only the scope supported by the evidence and record the next reassessment trigger.

illustrative diagram

An operating state can return to evaluation

An operating state can return to evaluationEvaluate: Record the population and criteria; insufficient evidence returns to observation. Trial: A bounded trial can expand when outcomes meet criteria or pause on a guardrail breach. Reassess: Material change or assessed remediation returns the system to evaluation.Population and criteriarecordedEvidence supports boundedtrialEvidence insufficientNew evidence collectedOutcomes meet agreedcriteriaGuardrail breachedRemediation assessedMaterial changeDefinedEvaluatedTrialObserveExpandedPaused
An operating state can return to evaluationEvaluate: Record the population and criteria; insufficient evidence returns to observation. Trial: A bounded trial can expand when outcomes meet criteria or pause on a guardrail breach. Reassess: Material change or assessed remediation returns the system to evaluation.Evidence supportsDefinedEvaluatedRecord populationand agreed criteriaInsufficient: ObserveNew evidence: evaluateBounded TrialMeets criteria:ExpandedGuardrail breach:PausedReturn to EvaluatedPaused: assess remedyExpanded:material change
Evaluate
Record the population and criteria; insufficient evidence returns to observation.
Trial
A bounded trial can expand when outcomes meet criteria or pause on a guardrail breach.
Reassess
Material change or assessed remediation returns the system to evaluation.

Representative structure from the design pack, not evidence of a measured deployment.

Filed under · Method · Method · Governance · Decisions Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.