Start a conversation Contact
← Research

The evaluation set is the asset. Build it before the system.

Teams treat evaluation data as something assembled to check a build. Reverse the order — the set is the durable artefact, and the system is the disposable one, because the model underneath it will be replaced within two years.

Every organisation building with models is accumulating two things. One is a system: prompts, retrieval, orchestration, an integration or two, all of it built against a model that will be retired on a published schedule. The other is a description of what good output looks like for this organisation, in this domain, for these users.

The first is what appears on the roadmap. The second is what has lasting value, and in most estates it does not exist as an artefact at all — it exists as opinions, distributed across the people who have looked at enough outputs to have them.

The claim: the evaluation set is the only part of an AI system that appreciates. The model will be replaced. The framework will be replaced. The set of examples that encodes what your organisation means by a correct answer survives all of it, and it is the only thing that makes a replacement decision answerable instead of a matter of taste.

What a usable set contains

Not a benchmark. A benchmark measures general capability against a public distribution. This measures whether a specific system does a specific job, and it looks nothing like a leaderboard.

What belongs in an evaluation set that is worth maintaining Fig. 01
  1. Part 01 Ordinary cases The bulk of real traffic, sampled rather than invented. Establishes whether the system does the everyday job.
  2. Part 02 Known-hard cases The ones experienced staff argue about. The set where a model change is most likely to show.
  3. Part 03 Failure cases Every incident, complaint and escalation, retained permanently. This is the part that compounds.
  4. Part 04 Refusal cases Questions the system should decline, or escalate. Otherwise nothing measures over-helpfulness.
  5. Part 05 Entitlement cases Requests from users who must not see certain material. The only way to test that retrieval respects permissions.

The third row is where the value accumulates. An organisation that adds every production failure to a permanent set, with the expected answer written down at the time, has after a year an asset no supplier can hand it and no benchmark can replace. An organisation that fixes each failure and moves on has a system that will regress on the same cases, repeatedly, and will find out from users.

The fifth row barely exists in practice and is the one that produces the worst incidents. Retrieval systems fail on permissions quietly: the answer looks correct, is correct, and was assembled from a document the person asking was not entitled to read. No general quality metric detects it. Only a test written by somebody who knew the permission boundary detects it.

Who writes the expected answers

This is the question that stalls the work, and the answer is not the AI team.

The expected answer is a statement about the domain — what a correct claims decision is, what a well-drafted clause looks like, which of two summaries a clinician would accept. The people who know that are the people currently doing the job, and the practical form of the work is a few hours of their time per month, on real cases, with disagreements recorded rather than resolved by whoever is most senior.

Recording disagreement is the part that gets skipped and the part that matters most. If two experienced reviewers disagree on a case, no model is going to be judged fairly on it, and the honest thing to do is mark it as contested and exclude it from the headline number while keeping it in the set. A set with no contested cases has usually been labelled by one person.

What it makes possible

Three decisions become answerable that are otherwise argued.

Whether to change model, when the provider retires the one you are on. That is not a hypothetical — model lifecycles are short and published deprecation calendars are the norm, as OpenAI’s own list of retirements shows, so the question arrives on a date somebody else chose. With a set, migration is a measurement. Without one, it is a rebuild followed by a period of hoping.

Whether a supplier’s system is better than yours. A supplier arriving with a leaderboard position is describing performance on a public dataset. Running their system against your set answers a question about your work — and the difference between those two things is the entire argument for owning the set.

Whether a change helped. Most prompt and retrieval changes are judged on a handful of examples someone tried by hand. A set of a few hundred, run automatically, turns that into a number, and the number occasionally says the change made things worse.

What this does not tell you

An evaluation set is not a safety case, and a system that scores well on one is not thereby suitable for a consequential decision. It measures the cases you thought to include, which is a strictly smaller thing than the cases that will arrive.

It also does not remove the need for monitoring in production. The set is fixed and the world is not. Its role is to catch regression, not novelty, and an organisation that stops watching live outputs because the offline number is healthy has swapped one blind spot for another.

The person who should act is whoever owns the budget for the next AI build. Fund the set separately from the system, give it an owner in the business, and insist that it is populated before the first model call goes to production. It is the cheapest thing on the plan and it is the only line item that will still be worth something after the model underneath it has been retired twice.

Filed under · Method · Evaluation · Method · Procurement Inference Institute · 02 Jul 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.