Request a scoping call Contact
← Research

Audit the benchmark before procuring against it

Annotation corrections changed agent rankings in two text-to-SQL benchmarks. Use public scores to inform a shortlist, then inspect the answer key and test candidates against cases the organisation can review.

Data / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A public leaderboard can help shortlist systems that generate queries from natural language. It is weaker evidence for the final procurement decision unless the buyer understands the questions, databases and reference answers behind the score.

The score compares system outputs with an annotated evaluation dataset. Its meaning depends on the accuracy of those annotations, the handling of ambiguous questions and the relevance of the test population. A named benchmark does not remove the need to inspect those conditions.

Audit the evaluation evidence before using a ranking to choose a supplier. The cited text-to-SQL study shows that correcting reference answers can change reported performance and agent order. It supports scrutiny of the answer key, rather than a conclusion that every benchmark or leading supplier is unreliable.

Somebody checked

Four researchers at the University of Illinois audited the annotations in two of the most heavily used text-to-SQL benchmarks and published the result in Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards in January 2026. They report annotation errors in 263 of the 498 examples in BIRD Mini-Dev, and in 76 of the 121 examples in Spider 2.0-Snow — 52.8 per cent and 62.8 per cent respectively. Their code and corrected data are public, which is the part that makes the finding arguable rather than merely alarming.

Error rates alone would be a quality complaint. The consequence is the second half of the work. The authors corrected a sample of one hundred examples from the BIRD development set and re-ran all sixteen open-source agents listed on the benchmark’s leaderboard against both versions, reporting relative performance changes from −7 per cent to 31 per cent and ranking movements of up to nine positions in either direction (Jin and colleagues, 2026).

One result in that paper deserves to be read twice. Rankings measured on the uncorrected sample tracked rankings on the full development set closely, at a Spearman correlation of 0.85. Rankings measured on the corrected sample did not, falling to 0.32 and losing significance (Jin and colleagues, 2026). The uncorrected benchmark was internally consistent. It agreed with itself, reproducibly, across a larger sample of the same flawed annotations.

A measurement that reproduces is not the same as a measurement that is right.

This is not an isolated finding about one dataset. The Agentic Benchmark Checklist, assembled in 2025 by researchers across several institutions, documents the same failure in different clothing across widely cited agent benchmarks — insufficient test cases, empty responses scored as successes — and estimates that such issues distort reported performance by as much as one hundred per cent in relative terms. The pattern is consistent. Benchmarks are built by people who need them to exist, under the same pressure as everyone else, and then they are used as though they were instruments.

Why the errors are invisible from where you are sitting

The useful part of the Illinois work, for a buyer, is the taxonomy rather than the headline rate. The authors sort annotation errors into four patterns, and the patterns differ in what a reviewer needs in order to see them at all.

One class is the one anybody can catch. Where the recorded query contradicts the question in plain terms — an inclusive range where the question asked for a strict inequality, a formula that computes something other than what was requested — the error is visible from the two texts alone. The largest class is not. Errors that come from a limited understanding of the schema or the data account for the majority of flawed examples in both benchmarks: a join on a non-unique key with no deduplication, a missing aggregation, a filter left out because a column was judged unimportant (Jin and colleagues, 2026). Those are only detectable by running queries against the actual database.

What a reviewer needs in order to see each kind of annotation error Fig. 01
  1. Level 01 The question and the recorded answer Catches an answer that contradicts what was asked. Needs only reading.
  2. Level 02 The schema Catches joins and aggregations that are wrong in structure.
  3. Level 03 The data Catches answers wrong because of what is in the tables. The largest class.
  4. Level 04 The domain Catches a label that encodes an incorrect fact about the subject.
  5. Level 05 A ruling on ambiguity Catches a question with more than one defensible answer, and needs an owner to settle it.

The review needs access to the material relevant to each error class. A question and reference query can reveal some contradictions, while schema, data and domain knowledge are needed for others. Ask which layers the buyer can inspect before treating the published score as a fully reviewable comparison.

What to do with a benchmark number in a procurement

Public benchmarks remain evidence about performance in their stated setting. Use them alongside a local evaluation whose schema, questions and adjudication the organisation can inspect. A buyer-controlled set also needs independent review: ownership does not make its labels correct.

The last of those is the one teams skip and the one that costs most later. Nearly a third of the flawed BIRD examples were flawed because the question itself permitted more than one reasonable reading (Jin and colleagues, 2026). An answer key built over your own data will inherit exactly that problem, and the resolution is not technical. Somebody has to decide what the business means by an active customer, and that decision has to be written down next to the answer, because it is part of the answer.

There is a regulatory edge to this as well. Article 15 of the EU AI Act requires that the levels of accuracy and the relevant accuracy metrics of a high-risk system be declared in the instructions for use. A number that began life as a leaderboard position can end up inside a declaration, at which point the provenance of the set it was measured on stops being an engineering preference. Whether any particular system falls in scope, and what a declaration has to say, is a question for counsel — this practice works on readiness and on the evidence that supports it, and does not certify anyone.

What this does not tell you

The audit covers two benchmarks in one task family. Text-to-SQL is unusually exposed, because the answer key is executable and therefore checkable at all — which is the reason these error rates are known rather than the reason they are high. For a summarisation or classification benchmark there is often no equivalent way to run the check, and the honest position is that the error rate is unmeasured rather than low.

Nor does any of this establish that the suppliers at the top of a leaderboard are the wrong choice. Correction moved rankings substantially in the study, and it moved some agents up. The finding is about the informativeness of the ordering, not about the direction of any particular error. A system may well be the right one and the number may still not be the reason.

The adjacent case for establishing a statistical baseline before the expensive model arrives concerns the comparison method. This article concerns the reference answers that make that comparison meaningful. Both need to be reviewable.

The evaluation section of a supplier assessment should identify the test population, label-review process and owner of ambiguous cases. Require candidates to run against that evidence under matched conditions. The procurement decision will then have a basis that can be revisited when the public benchmark changes.

Filed under · Data · Evaluation · Benchmarks · Provenance Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.