Evaluate the system where its decisions are made
A model score tells a buyer too little about a live AI workflow. Evaluate the inputs, tools, handoffs and outcomes that create the decision, and specify what evidence would stop deployment.
A supplier presents a benchmark score for an assistant that will help staff review customer cases. The benchmark asks the model to answer isolated questions. The live service will retrieve case material, call a tool, draft a recommendation and hand it to a reviewer. A buyer can accept the score and still have no evidence about the decision path they intend to put into use.
The unit of evaluation should be the system in its operating context. Name the decision and run the workflow through situations that make it difficult: missing records, conflicting sources, an unavailable tool, a change in permissions and a reviewer who disagrees. A model benchmark can remain one instrument, but it cannot stand in for the whole acceptance decision.
Turn a use case into an event
NIST’s TEVV-Athlon initial public draft offers a way to construct assessments around organisational objectives. It describes events and tools producing data on measurement concepts of interest. It is a draft open for comment through 6 October 2026, not a final standard or a certification route. Its useful move is to start from the effect the organisation needs to understand, then choose the event and instrument that can observe it.
For the case-review assistant, an event might start when a reviewer receives an incomplete file and end when they approve, amend or reject a recommendation. The team can record the retrieved documents, permission checks, tool results, model output, reviewer intervention and eventual customer outcome. This is more demanding than counting answer matches, because the failure may occur between components rather than in the model response.
The FCA’s account of AI Live Testing makes the same system-level distinction for its financial-services programme. It considers the model, deployment context, governance, human involvement, evaluation and input and output controls together. That is a description of the FCA programme, not an assurance that any particular firm’s system is safe.
Write the acceptance rule before collecting a score
Choose cases that reflect the service’s actual boundary, including cases the system should refuse or escalate. Keep an evaluation set with known versions and provenance. The existing evaluation-set article explains why that set is an asset. The next task is to embed it in a system event, so the test can catch a correct model answer delivered to the wrong person or an accurate recommendation based on an unauthorised document.
For each event, record a baseline workflow, a primary outcome, guardrails and a stop rule. A time saving has little value if reviewers miss more harmful cases. A good answer has little value if the retrieval layer exposes restricted records. The team should also specify who adjudicates disagreement, when a sample is large enough to support a decision and what has to be retested after a model, tool or policy change.
Use both observations and traces. The observation says what happened to the work and the person affected. The trace helps explain which part of the system created it. Neither alone is enough: a detailed trace without an outcome is an engineering record, while an outcome without a trace may not tell the owner what to repair.
- Baseline
- A sample of cases processed under the existing reviewer workflow, with outcomes and review time recorded.
- Outcome
- Correct, timely case decisions after retrieval, tool use and human review, measured on comparable case types.
- Guardrails
- Restricted-document exposure, unsupported recommendations, missed escalation and reviewer workload.
- Decision rule
- Release only for case types whose outcome improves without crossing a guardrail. Stop or redesign when a harmful case cannot be traced and corrected.
For example, include a case where the assistant retrieves a current policy and an older, conflicting one. A high answer score may hide that the system cited the wrong version. The event test should show which document was retrieved, what the assistant said, whether the reviewer caught the conflict and what the customer eventually received. Repeat with the older policy removed from the user’s permissions. The system should now refuse to use it, even if the model can still write a plausible answer.
Assign an adjudicator before running the set. When reviewers disagree about the right outcome, record the dispute rather than forcing a clean label into the score. Those disagreements reveal whether the policy itself needs clarification. They also prevent the team from quietly changing its measure after seeing the model’s results. Keep the event, trace and adjudication together so a later tool update can be tested against the same decision path.
What this does not tell you
NIST’s draft supplies a framework for designing assessment, not a universal list of metrics or a pass mark. The FCA programme is a particular regulatory setting. Neither establishes that a benchmark is useless or that every system needs a live trial. The appropriate evidence depends on the decision, harms, users and operating environment.
The architecture owner should be able to hand a review board an event, the instruments that observed it, the results and the rule that turns those results into a proceed, redesign or stop decision. Without that chain, a high score is merely a property of a test the buyer did not actually mean to buy.