THE SIGNAL IN ONE SENTENCE
Language models used to grade other models may vary enough across repeated requests that tiny leaderboard differences cannot be treated as stable measurements.
01
WHAT ACTUALLY CHANGED
Researchers audited language models used as judges in two preregistered campaigns. These judges increasingly filter training data, compare generated answers, and determine leaderboard positions, so the study treated each endpoint as a measurement instrument rather than assuming its model name guaranteed stable behavior.
Across 52,988 request attempts, rankings repeated within the same testing window reached a Spearman agreement of 0.400 against a required threshold of 0.90. Byte-identical requests replayed the next day reached 0.78 agreement against a required 0.99.
The researchers tested four providers and found median reliability between 0.74 and 0.88. Switching metrics or changing sampling settings did not repair the problem on the tested grid. Self-hosting helped only while the serving system remained quiet.
Several candidate differences were seven orders of magnitude smaller than the judge's measured noise floor. The authors estimate that a pilot using roughly two percent of the eventual request volume would have shown that the planned reliability gates were unreachable before the full campaigns began.
02
WHY THIS MATTERS
A leaderboard can look precise while recording noise. If identical inputs produce different rankings, a fraction-of-a-point victory may say more about the serving endpoint than the models being compared.
The problem reaches beyond public benchmarks. AI judges are used to select training examples, evaluate safety behavior, compare prompts, and approve deployment changes. An unstable judge can quietly turn a quality-control system into a random policy generator wearing a lab coat.
The practical response is not to abandon automated evaluation. It is to measure the evaluator first, report uncertainty, use repeated samples, and refuse to treat differences below the instrument's noise floor as meaningful.
03
WHERE IT COULD HELP
- Run repeatability checks before trusting an AI-scored benchmark
- Measure judge variance before defining release gates
- Compare shared endpoints with pinned or self-hosted snapshots
- Report confidence intervals instead of suspiciously precise rankings
KEEP A HAND ON THE WHEEL
This is a preregistered preprint, not peer-reviewed consensus. Its results concern the tested prompts, providers, and shared serving conditions. Independent replication should test other judges and tasks, but teams do not need to wait before adding a small reliability pilot to their own evaluation process.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
Benchmark
A fixed test used to compare how systems perform on the same tasks.
OPEN GLOSSARY CARD
Model snapshot
A fixed, dated version of a model whose behavior can be tested and referenced consistently.
OPEN GLOSSARY CARD
Evaluation gate
A required test or review that a system must pass before it advances to the next stage.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 4, 2026.
PUBLICATION RECEIPT: Revision 1. Approved by Zak and published September 4, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US