Request a scoping call Contact
← Research

Build evaluation evidence before approving the system

Fund and maintain evaluation evidence alongside the AI service. Representative cases, reviewed expectations and explicit access boundaries make supplier, model and workflow changes easier to assess.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

An AI service needs both an implementation and a documented account of acceptable work. Prompts, retrieval, models and integrations can change. The organisation’s evaluation evidence provides a reference for assessing those changes against its intended users, tasks and consequences.

Make that reference an owned asset rather than a collection of informal opinions. Record examples, expected outcomes, reviewer judgements and unresolved disagreements. Maintain it with the service so a future team can understand which claims the evaluation supports.

Invest in the evaluation set before accepting the service, and maintain it as the task changes. It can retain useful knowledge across model replacements. Its value is not automatic: stale cases, leaked hold-outs or obsolete expectations can make a previously useful set misleading.

What a usable set contains

A public benchmark measures the tasks and population it contains. A local evaluation asks whether the proposed service can complete this organisation’s work. Define that use boundary before choosing cases. Public and local evidence can complement each other, but they answer different questions.

What belongs in an evaluation set that is worth maintaining Fig. 01
  1. Part 01 Ordinary cases The bulk of real traffic, sampled rather than invented. Establishes whether the system does the everyday job.
  2. Part 02 Known-hard cases The ones experienced staff argue about. The set where a model change is most likely to show.
  3. Part 03 Failure cases Reviewed failure cases, retained under appropriate data and retention rules.
  4. Part 04 Refusal cases Questions the system should decline, or escalate. Otherwise nothing measures over-helpfulness.
  5. Part 05 Entitlement cases Requests from users who must not see certain material. Test retrieval entitlement separately from answer quality.

Add reviewed failure cases where retention is appropriate. Preserve the cause, expected behaviour and relevant context so the next release can be tested against the same problem. Use lawful retention and suitable de-identification rather than keeping every incident record permanently. A recurring-case set helps detect regressions without representing all future failures.

Include tests of entitlement independently of answer quality. A correct answer from a restricted document can still breach the service’s access rule. Tests need allowed and denied users, changed permissions and cached responses. Quality measures alone do not establish that those boundaries are enforced.

Who writes the expected answers

Domain owners should define and review the expected outcomes with evaluation specialists. The AI team contributes the test design and execution. Neither technical ownership nor seniority alone settles a disputed domain judgement.

A claims outcome, contractual clause or clinical summary needs a reviewer qualified for that context. Budget the review work according to case volume and consequence. Record the criteria and source evidence so later reviewers can distinguish a legitimate disagreement from an incorrect label.

Mark contested cases and retain the competing interpretations. Agree adjudication and scoring rules before comparing models. Some cases may need exclusion from a single-answer metric, while others support a range of acceptable outcomes. Report their count and treatment rather than silently manufacturing certainty.

What it makes possible

Three decisions become answerable that are otherwise argued.

Model replacement. OpenAI’s published deprecations illustrate that hosted models can retire on a provider’s schedule. A maintained local set helps evaluate a candidate replacement under the actual service conditions. It does not remove migration work or establish that an unavailable model can be restored.

Whether a supplier’s system is better than yours. A supplier arriving with a leaderboard position is describing performance on a public dataset. Running their system against your set answers a question about your work — and the difference between those two things is the entire argument for owning the set.

Service changes. Run prompt and retrieval changes against the relevant cases and compare both outcomes and uncertainty. Set size follows the questions the comparison must answer. Keep a protected hold-out where needed, and use production monitoring to detect changes outside the offline set.

What this does not tell you

An evaluation set is not a safety case, and a system that scores well on one is not thereby suitable for a consequential decision. It measures the cases you thought to include, which is a strictly smaller thing than the cases that will arrive.

It also does not remove the need for monitoring in production. The set is fixed and the world is not. Its role is to catch regression, not novelty, and an organisation that stops watching live outputs because the offline number is healthy has swapped one blind spot for another.

The budget owner should fund the evaluation separately, name its business owner and require evidence before production approval. Record how it will be maintained and which uses it cannot validate. That retained knowledge supports later procurement and architecture decisions without treating the current model as permanent.

illustrative diagram

Define the target population before comparing

Define the target population before comparingPopulation: The fictional new-account rollout holds out whole accounts. Comparison: Alder and Birch train; Cedar and Dune test against the incumbent on the same cases. Decision: Agreed thresholds permit bounded expansion, further evidence or a narrower scope.YesUncertainNoNew-account rolloutHold out whole accountsTrain on Alder and BirchEvaluate on Cedar and DuneCompare with incumbent onsame casesMeets agreed thresholds?Bounded expansionGather more evidenceRetain supported scope
Define the target population before comparingPopulation: The fictional new-account rollout holds out whole accounts. Comparison: Alder and Birch train; Cedar and Dune test against the incumbent on the same cases. Decision: Agreed thresholds permit bounded expansion, further evidence or a narrower scope.New-account rolloutHold out wholeaccountsTraining: Alder, BirchTest: Cedar, DuneSame-case comparisonwith the incumbentAgreed thresholdsYes: boundedexpansionUncertain:more evidenceNo: supported scope
Population
The fictional new-account rollout holds out whole accounts.
Comparison
Alder and Birch train; Cedar and Dune test against the incumbent on the same cases.
Decision
Agreed thresholds permit bounded expansion, further evidence or a narrower scope.

Representative structure from the design pack, not evidence of a measured deployment.

Filed under · Method · Evaluation · Method · Procurement Inference Institute · 02 Oct 2026 (updated)

Related engagement

The decision behind this article

A clear design your team or chosen delivery partner can build from.

Explore AI Architecture →

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.