An automated judge is an instrument. Calibrate it or do not read it.
An automated judge can produce consistent scores while giving biased comparisons. Validate it against human judgements on the intended task, test ordering effects and recheck the instrument when the rubric or model changes.
Keep development and test groups separate before comparing performance. No measured results are shown.
An evaluation score can become a release gate before the team has validated the model producing it. When another language model judges each answer, the gate depends on that judge’s rubric, behaviour and agreement with informed human reviewers.
Automated judging can increase evaluation coverage and support frequent releases. Its value depends on testing the judge as part of the measurement system. Repeatable output is useful evidence about stability, but does not establish that the score measures the quality the organisation intended.
Validate an automated judge before using its scores for consequential comparisons. Check agreement with appropriate human reviewers, sensitivity to answer order and the failure modes omitted by the rubric. Preserve the judge version and review the instrument when its inputs or purpose change.
Consistency is not accuracy
The natural way to check a judge is to run it twice and see whether it agrees with itself. It usually does, and that is the trap. A large study across twenty-one judge models and several hundred thousand individual judgements reported exactly this pattern — very high test-retest reliability coexisting with severe positional bias in judges already deployed in production. The paper, Reliability without Validity, is worth reading in full if your release gate depends on one of these.
Positional bias is the best-studied of the family: presented with two candidate answers, judges tend to prefer one position over the other regardless of content. It has its own literature — the systematic study of position bias in LLM-as-a-judge sets out how it is measured — and it is joined by preferences for longer answers, for answers written in a familiar style, and for answers produced by the judge’s own model family.
A repeatability check can reproduce a systematic preference each time. That is why a stable score needs an independent validity check. The relevant question is whether the instrument discriminates between outputs for the task, rather than merely agreeing with its previous answer.
What the errors do to a decision
| What the harness reports | What may actually be true |
|---|---|
| Variant B scored higher than variant A | Variant B was shown second, and this judge prefers the second answer |
| Quality improved after the prompt change | The change made answers longer, and this judge rewards length |
| Our own model outperformed the alternative | The judge shares a family with our model and prefers its phrasing |
| The score held steady across the release | The failure mode that appeared is one the rubric never asked about |
The last row is the one that costs most and gets discussed least. A judge scores what the rubric names. Anything the rubric does not name — a tone that is wrong for the audience, a citation to a document the user cannot access, a confident answer to a question that should have been refused — is not scored as bad. It is not scored at all, and the aggregate is unchanged.
What calibration actually involves
Treat judge validation as maintained evaluation work. Its scope depends on the decision and the range of outputs, and it must be repeated when a material change can affect the comparison.
Build a reference set covering the intended domain and consequential failure modes. Use clear human review criteria, record disagreement and report uncertainty. There is no universal sample size: the set must support the distinctions the release decision requires, including important segments.
Run pairwise comparisons in both orders to expose position sensitivity. Record disagreements and use a stated aggregation or escalation rule. Swapping positions is a diagnostic control, rather than proof that all ordering bias has been removed.
A change in judge version can alter the meaning of a historical score. Pin the version where possible and retain a common reference set for comparison across changes. Revalidate rather than assuming that old and new scores remain directly comparable.
What this does not tell you
Automated judging can remain valuable at volume when its validity is monitored. Combine it with targeted human review and deterministic checks where those can test the requirement more directly. Match the effort to the consequence of an incorrect release decision.
It is also not a claim that any particular published agreement figure transfers to your task. It does not. Agreement is a property of the rubric, the domain and the population of outputs, and a judge that tracks human preference well on general assistant responses may be close to useless on clinical summaries or contract clauses.
The owner of the release gate should be able to explain how the judge agrees with informed reviewers on the organisation’s data, which biases were tested and when validation last ran. Where that evidence is missing, restrict the score to exploratory use until the gate has a defensible measurement basis.