THE SIGNAL IN ONE SENTENCE

Language models used to grade other models may vary enough across repeated requests that tiny leaderboard differences cannot be treated as stable measurements.

01

WHAT ACTUALLY CHANGED

Researchers audited language models used as judges in two preregistered campaigns. These judges increasingly filter training data, compare generated answers, and determine leaderboard positions, so the study treated each endpoint as a measurement instrument rather than assuming its model name guaranteed stable behavior.

Across 52,988 request attempts, rankings repeated within the same testing window reached a Spearman agreement of 0.400 against a required threshold of 0.90. Byte-identical requests replayed the next day reached 0.78 agreement against a required 0.99.

The researchers tested four providers and found median reliability between 0.74 and 0.88. Switching metrics or changing sampling settings did not repair the problem on the tested grid. Self-hosting helped only while the serving system remained quiet.

Several candidate differences were seven orders of magnitude smaller than the judge's measured noise floor. The authors estimate that a pilot using roughly two percent of the eventual request volume would have shown that the planned reliability gates were unreachable before the full campaigns began.

02

WHY THIS MATTERS

A leaderboard can look precise while recording noise. If identical inputs produce different rankings, a fraction-of-a-point victory may say more about the serving endpoint than the models being compared.

The problem reaches beyond public benchmarks. AI judges are used to select training examples, evaluate safety behavior, compare prompts, and approve deployment changes. An unstable judge can quietly turn a quality-control system into a random policy generator wearing a lab coat.

The practical response is not to abandon automated evaluation. It is to measure the evaluator first, report uncertainty, use repeated samples, and refuse to treat differences below the instrument's noise floor as meaningful.

FIG. 037MEASURE THE RULER FIRST
1IDENTICAL INPUT→
2REPEAT JUDGE→
3MEASURE VARIANCE→
4SET NOISE FLOOR→
5TRUST LARGE GAPS
Before a judge decides which model won, repeated trials establish whether the judge can reproduce its own ruling.

03

WHERE IT COULD HELP

  • Run repeatability checks before trusting an AI-scored benchmark
  • Measure judge variance before defining release gates
  • Compare shared endpoints with pinned or self-hosted snapshots
  • Report confidence intervals instead of suspiciously precise rankings

KEEP A HAND ON THE WHEEL

This is a preregistered preprint, not peer-reviewed consensus. Its results concern the tested prompts, providers, and shared serving conditions. Independent replication should test other judges and tasks, but teams do not need to wait before adding a small reliability pilot to their own evaluation process.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on September 4, 2026.

PUBLICATION RECEIPT: Revision 1. Approved by Zak and published September 4, 2026.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US