Request a scoping call Contact
← Research

Only the questions where they disagree are deciding your model choice

A paired model comparison needs more than two average scores. Record item-level disagreements, uncertainty and the smallest difference worth acting on, then assess whether the evaluation can resolve that decision.

Method / Conceptual study
Separate before comparing.
  1. Development groups
  2. Hold-out boundary
  3. Test groups

Keep development and test groups separate before comparing performance. No measured results are shown.

A model-selection table may show one candidate a few points above another on the organisation’s questions. Before accepting the ordering, review the number and independence of items, how answers were marked and the uncertainty around the difference. The available marking budget alone does not establish that the set can resolve the choice.

Local questions can make an evaluation relevant to the intended task, but their provenance and labels still need review. The earlier article on benchmark auditing addresses that evidence quality. The separate question here is whether the comparison has enough statistical resolution for the decision.

For paired binary correctness scores, the difference comes from items on which the candidates receive different marks. Shared successes and failures still describe absolute performance, but cancel in the paired difference. Record the disagreements and their direction, then calculate uncertainty using the relevant sampling assumptions.

What a single score is reporting

A score is not a measurement of a system. It is an estimate, drawn from a sample of the questions the organisation could have asked, of how the system would do across all of them. That is the framing Evan Miller sets out in Adding Error Bars to Evals, which treats an evaluation as an experiment and analyses it as one.

The arithmetic follows from the framing. For a score expressed as a proportion, the standard error of the mean is the square root of p(1 − p)/n. Take a set of fifty questions and a system answering four in five correctly, which are stated assumptions rather than measurements: the standard error is a little under six points, so the interval around that score spans roughly eleven points either side (Miller, 2024). Two hundred questions brings it to about five and a half points. A thousand brings it to about two and a half.

Under the illustrative independent-score assumptions above, a conventional approximate comparison needs a gap of around sixteen percentage points to reach the stated significance level. This is not a universal decision threshold. Pairing, clustering, sample size and the selected test change the requirement.

Pairing is free, and it shows you where the evidence is

If both candidates answered the same questions, the summary scores are the wrong thing to compare. The right comparison is item by item, because the two systems find the same questions hard and the same questions easy, and analysing the paired differences removes that shared difficulty from the noise. Miller lists this as one of his recommendations, alongside computing standard errors at all, clustering them when questions come in related groups, and using power analysis to establish whether a set is capable of testing the hypothesis it is being handed (Miller, 2024).

Item-level pairing can improve precision without collecting a separate set for each candidate. If two systems receive the same binary mark on ninety of one hundred items, ten disagreements drive the paired difference. The shared items still matter when assessing whether either candidate meets an absolute quality requirement.

Two systems that agree on nine questions in ten have handed you a ten-item experiment, whatever the headline number says.

Ten items is a small experiment and the threshold it imposes is severe. Under a two-sided test at the conventional level, at least nine of those ten disagreements have to fall the same way before the result clears it. The binomial arithmetic is worth doing once: from ten fair tosses, a nine-one split or better arrives about one time in fifty, and an eight-two split or better above one time in ten. A seven-three split, which looks decisive written on a slide, is roughly what a coin produces on one afternoon in three.

Somebody has measured how often this goes wrong

This is not a theoretical objection, and it is not confined to sets built in a fortnight. In Resolution Diagnostics for Paired LLM Evaluation, published on 28 May 2026, Anany Kotawala frames paired evaluation as hypothesis testing and defines a resolution ratio: the number of items an evaluation actually has, divided by the number it would need. Applied to published leaderboards, the author reports 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9 adjacent-rank pairs in the MMLU-Pro top ten as unresolved at a significance level of 0.05 and 80 per cent power, and reports that the pattern survives correction for multiplicity and sequential testing (Kotawala, 2026).

The same paper carries a finding that ought to travel further than it has. The common shortcut for sizing these comparisons — a Cohen-h effect size with a correlation adjustment — deviates from the correct requirement by roughly a factor of two in close comparisons, and three of five standard calculators the author examined carry that error (Kotawala, 2026). A team that did the statistical work rather than skipping it can still have planned a set half the size it needed.

Large leaderboards can remain unable to resolve close pairs. A smaller local set may resolve a larger difference or fail to resolve a close one. Calculate the requirement for the local margin and disagreement rate rather than infer sufficiency from the size of another benchmark.

Sizing the set before the comparison is run

The order matters more than the formulae. The number of items is not a budget decision to be reconciled with statistics afterwards — it is derived from the difference the organisation needs to be able to see.

The order in which the size of an evaluation set is decided Fig. 01
  1. 01 The margin The smallest difference that would change which system you choose.
  2. 02 The disagreement How often the two candidates answer the same item differently.
  3. 03 The count Items required to resolve that margin, at a stated confidence and power.
  4. 04 The comparison Run once, on that set, reported with an interval.

Worked through with stated assumptions, using the method in Miller: suppose the difference worth acting on is five points, the two candidates disagree on about one item in five, and the test is run at a significance level of 0.05 with 80 per cent power. The comparison needs somewhere in the order of a hundred and twenty disagreements to resolve, and at that rate of disagreement that means a set in the order of six hundred questions. Change the assumptions and the number moves a long way — a margin of ten points is far cheaper to see than a margin of two. The discipline is doing the calculation before the marking starts rather than after the table is drawn.

Ask the supplier for item-level marks, disagreement counts and the uncertainty calculation. Check adjudication and sampling as well as arithmetic. Having these records makes the comparison inspectable, but does not by itself establish that the experiment is valid.

What this does not tell you

Statistical power does not make an evaluation representative. It describes the chance of detecting a specified difference under stated assumptions. A precise result on the wrong tasks remains a weak basis for deployment, so assess population, independence and label quality alongside resolution.

Nor does it address whether the answers are right. That is a separate and compounding problem — correcting labels has been shown to move rankings on sets everybody trusted — and neither piece of work substitutes for the other.

A resolved difference is also not a claim about production. It establishes that one system answered a fixed set of questions better than another. It does not establish that either behaves as required once it is in front of people, and no evaluation this practice designs certifies anyone or anything. What an evaluation can do is make a decision defensible on the evidence it actually contains, which is a narrower and more useful thing.

Before the next comparison, agree the smallest meaningful difference and estimate the evidence needed to resolve it. If the required set is impractical, state that limitation. The selection may then depend on other justified criteria, including cost, operating effort and switching constraints, rather than an unsupported ranking.

Filed under · Method · Evaluation · Method · Procurement Inference Institute · 02 Oct 2026 (updated)

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.