PhaseSuperseded
Phases 1–2 — prototype and instrument repair
The historical v1 prototype and the mechanical re-adjudication layer built after it exposed analyst discretion.
Postmortemv1 paper (frozen)Metrics spec
Phase 1 was the algorithmic-strategy prototype: 7 classic games, 8 strategies, 40 experiments, and a 2,149-word v1 paper whose 11 claims were all called “supported.” Mechanical re-adjudication in Phase 2 revised that record to 6 supported, 1 refuted, 4 inconclusive — the refuted claim (TFT >50% cooperation vs Always Defect; observed 0.02) became dead prediction #1. Four process errors were documented: a literature transplant, payoff totals presented as per-round, metrics applied to undefined game classes, and single unseeded runs treated as facts.
Phase 2 rebuilt the instrument: seeded mulberry32 randomness, 20-seed replicate batches, a versioned per-class metric suite, predicate-based mechanical adjudication (verdicts never hand-set; predicates immutable after first adjudication), and a 396-run backfill with zero drift. Neither phase carries confirmatory weight; they exist in the record because the correction is part of the result.