Only the questions where they disagree are deciding your model choice
A comparison between two systems on the same evaluation set is carried entirely by the items the two answer differently, and that count is usually in the tens. Recent work reports that many published pairwise rankings are not resolved at conventional significance and power.
The table has two rows in it. One system, then the other, one score each, and the second is a few points above the first. Somebody has assembled questions from the organisation’s own material, run both candidates against them and marked the answers, and the table is the return on that work. The recommendation follows the arrow. Nobody asks how many questions were in the set, because its size was settled by how much marking one person could do in a fortnight.
Using your own questions rather than a public leaderboard is the right decision, and we have argued for it — a benchmark is a dataset nobody audited. Owning the set does not make its verdict readable. The size that makes a set affordable to build is very often the size that makes it unable to answer the question it was built for.
Here is the claim. A comparison between two systems on the same questions rests entirely on the questions the two answer differently. Every item both get right and every item both get wrong contributes nothing at all to the difference between them. On a set of the size enterprises actually build, the number of items carrying the comparison is usually in the tens, and a gap of a few points across them is not distinguishable from the toss of a coin.
What a single score is reporting
A score is not a measurement of a system. It is an estimate, drawn from a sample of the questions the organisation could have asked, of how the system would do across all of them. That is the framing Evan Miller sets out in Adding Error Bars to Evals, which treats an evaluation as an experiment and analyses it as one.
The arithmetic follows from the framing. For a score expressed as a proportion, the standard error of the mean is the square root of p(1 − p)/n. Take a set of fifty questions and a system answering four in five correctly, which are stated assumptions rather than measurements: the standard error is a little under six points, so the interval around that score spans roughly eleven points either side (Miller, 2024). Two hundred questions brings it to about five and a half points. A thousand brings it to about two and a half.
Two systems scored independently on fifty questions each therefore need a gap of around sixteen points before the ordering means anything at all. Most enterprise comparisons are decided on three or four.
Pairing is free, and it shows you where the evidence is
If both candidates answered the same questions, the summary scores are the wrong thing to compare. The right comparison is item by item, because the two systems find the same questions hard and the same questions easy, and analysing the paired differences removes that shared difficulty from the noise. Miller lists this as one of his recommendations, alongside computing standard errors at all, clustering them when questions come in related groups, and using power analysis to establish whether a set is capable of testing the hypothesis it is being handed (Miller, 2024).
Pairing buys precision for nothing. It also makes visible how thin the evidence usually is, which is the part worth sitting with, because agreement carries no information about which system is better. If two candidates agree on ninety questions out of a hundred, the entire comparison lives in the remaining ten.
Two systems that agree on nine questions in ten have handed you a ten-item experiment, whatever the headline number says.
Ten items is a small experiment and the threshold it imposes is severe. Under a two-sided test at the conventional level, at least nine of those ten disagreements have to fall the same way before the result clears it. The binomial arithmetic is worth doing once: from ten fair tosses, a nine-one split or better arrives about one time in fifty, and an eight-two split or better above one time in ten. A seven-three split, which looks decisive written on a slide, is roughly what a coin produces on one afternoon in three.
Somebody has measured how often this goes wrong
This is not a theoretical objection, and it is not confined to sets built in a fortnight. In Resolution Diagnostics for Paired LLM Evaluation, published on 28 May 2026, Anany Kotawala frames paired evaluation as hypothesis testing and defines a resolution ratio: the number of items an evaluation actually has, divided by the number it would need. Applied to published leaderboards, the author reports 11 of 40 Open LLM Leaderboard v1 pairwise comparisons and 4 of 9 adjacent-rank pairs in the MMLU-Pro top ten as unresolved at a significance level of 0.05 and 80 per cent power, and reports that the pattern survives correction for multiplicity and sequential testing (Kotawala, 2026).
The same paper carries a finding that ought to travel further than it has. The common shortcut for sizing these comparisons — a Cohen-h effect size with a correlation adjustment — deviates from the correct requirement by roughly a factor of two in close comparisons, and three of five standard calculators the author examined carry that error (Kotawala, 2026). A team that did the statistical work rather than skipping it can still have planned a set half the size it needed.
Read that against your own position. Those leaderboard sets run to thousands of items. If sets that large cannot separate systems standing next to each other in a ranking, two hundred questions will not separate two suppliers whose demonstrations both went well.
Sizing the set before the comparison is run
The order matters more than the formulae. The number of items is not a budget decision to be reconciled with statistics afterwards — it is derived from the difference the organisation needs to be able to see.
- 01 The margin The smallest difference that would change which system you choose.
- 02 The disagreement How often the two candidates answer the same item differently.
- 03 The count Items required to resolve that margin, at a stated confidence and power.
- 04 The comparison Run once, on that set, reported with an interval.
Worked through with stated assumptions, using the method in Miller: suppose the difference worth acting on is five points, the two candidates disagree on about one item in five, and the test is run at a significance level of 0.05 with 80 per cent power. The comparison needs somewhere in the order of a hundred and twenty disagreements to resolve, and at that rate of disagreement that means a set in the order of six hundred questions. Change the assumptions and the number moves a long way — a margin of ten points is far cheaper to see than a margin of two. The discipline is doing the calculation before the marking starts rather than after the table is drawn.
Ask a supplier presenting a comparison for the third and fourth of those. The request is unusual enough to be informative on its own, and a supplier who has the numbers to hand has run a real experiment.
What this does not tell you
None of this makes a set representative. Power is a statement about noise and nothing else: it establishes that a difference is unlikely to be an accident of which questions were drawn, and it says nothing whatever about whether those questions resemble the work. A set of six hundred items sampled from the wrong part of the estate resolves the wrong difference precisely.
Nor does it address whether the answers are right. That is a separate and compounding problem — correcting labels has been shown to move rankings on sets everybody trusted — and neither piece of work substitutes for the other.
A resolved difference is also not a claim about production. It establishes that one system answered a fixed set of questions better than another. It does not establish that either behaves as required once it is in front of people, and no evaluation this practice designs certifies anyone or anything. What an evaluation can do is make a decision defensible on the evidence it actually contains, which is a narrower and more useful thing.
The person this changes is whoever is about to sign off a model selection. Before the next comparison runs, write down the smallest difference that would change the answer, and work out how many items it would take to see it. If the number comes back larger than the set anybody is willing to build, that is not a failure of the exercise. It is the finding: the evaluation was never going to settle this, and the choice has to be made on the grounds that remain — the cost of running each system, the effort of operating it, and how hard it would be to leave.