Write down what would make you stop
Agree the evidence that would pause, narrow or end an AI use before approving it. Name the measurement, threshold and accountable owner, then connect review findings to a deliberate operating decision.
Keep development and test groups separate before comparing performance. No measured results are shown.
An AI review needs evidence against the purpose for which the work was funded. Adoption, released features and favourable feedback can be useful observations without establishing the promised outcome. Agree the conditions for reconsideration before relying on those observations to justify continued spending.
If the acceptance process cannot produce a pause, restriction or rejection, its decision scope is incomplete. State the findings that would lead to each response. That makes an unfavourable result usable instead of leaving the review team to negotiate its meaning afterwards.
Set measurable reconsideration conditions before the result is known. They should identify the population, threshold, review time and person authorised to act. Conditions may later change, but the reason and evidence for that change must be recorded.
What a stopping condition is, and what it is not
It is not a risk register entry. It is not a phrase about monitoring. It is a sentence of the form: if this measurement, on this population, is worse than this value, we stop — and a named person who is accountable for acting on it.
Three properties make it real.
It is measurable with something that already exists, or that the project builds first. A condition that depends on data nobody collects is a condition that will never trigger.
It is checked on a schedule that is set in advance. Conditions checked when somebody remembers are checked when things are going well.
Name the immediate response to a breach: pause, restrict, roll back, escalate or enter a defined review state. A review can be appropriate if its authority, deadline and interim controls are explicit. An unspecified promise to review does not establish how the service behaves while the issue is unresolved.
None of this is novel as a principle. The Manage function of the NIST AI Risk Management Framework asks organisations to decide, on evidence, whether a system continues in service or is decommissioned. What is missing in practice is not the principle. It is the sentence that would let the decision be made by anyone other than the team whose work is being decided about.
Under what measured result does this system stop, narrow or roll back?
- Accuracy on the held-back set falls below the threshold the process was designed around. Requires an evaluation set that exists before launch and is not used for tuning.
- Errors fall unevenly across groups, or a single failure exceeds a stated severity. Needs the outcome data disaggregated. If it is not collected, this condition cannot fire.
- The measured benefit does not appear by the date it was forecast for. Assess the measured benefit against its agreed value and time conditions.
Choose the scope the evidence supports
Swipe or scroll for the full diagram →
What decision does the measured result support?
Owner: Accountable operating owner. Threshold: Record the quality, harm and value thresholds, target population and review date before collecting the result.
- Expand
- Expand only when the comparison clears every agreed threshold within its evaluated population.
- Retain narrower scope
- Retain the supported population or return to the incumbent when a guardrail fails.
- Gather more evidence
- Gather more evidence when the sample or outcome window cannot establish a decision.
Representative decision structure; the operating owner sets local thresholds, not this diagram.
Evidence and named thresholds lead to expansion, narrower scope or further observation.
Reviewed 2026-10-02
Include the value condition
Quality, harm and value require separate decisions. A service may perform accurately while failing to deliver an affordable benefit, or deliver a benefit while breaching a guardrail. Make those conditions visible to the accountable sponsor rather than compressing them into one success score.
The sponsor and relevant operating owners should agree the conditions at funding and acceptance. Include the evaluation team’s judgement about what can be measured. Preserve the approval record so later review can compare the outcome with the original assumptions.
Applying these conditions across a portfolio can identify work that should continue, narrow or stop. The result depends on the evidence for each use. It does not establish that most AI programmes are unevaluated or that a particular share should be cancelled.
What it does for the systems that survive
The argument for stopping conditions is usually made as a risk argument. The stronger argument is a commercial one.
A useful decision record connects the baseline, measurements, thresholds and owner. The same evidence can support expansion when outcomes meet the criteria. Merely writing a stopping condition does not establish that the data exists or that the service meets it, so test the review process as well.
What this does not tell you
A stopping condition does not make a system safe, and it is not a substitute for an impact assessment where one is required. It is the mechanism that gives an assessment somewhere to land — a way for a finding to have a consequence other than a paragraph.
A breach requires the agreed response and an accountable decision. Some conditions demand immediate suspension. Others permit a constrained review. Any exception needs authority, a documented reason, an expiry and interim controls. Recording an override does not by itself make it safe or lawful.
The funding owner should require the conditions for continuation, restriction and withdrawal with the initial proposal. Confirm that each can be measured and acted on. At review, accept only the scope supported by the evidence and record the next reassessment trigger.
An operating state can return to evaluation
- Evaluate
- Record the population and criteria; insufficient evidence returns to observation.
- Trial
- A bounded trial can expand when outcomes meet criteria or pause on a guardrail breach.
- Reassess
- Material change or assessed remediation returns the system to evaluation.
Representative structure from the design pack, not evidence of a measured deployment.