Program
The five-phase experimental program
Sequential, sealed, and mechanically adjudicated — with each phase's question set by the previous phase's failure.
| Phase | Status | Question and answer |
|---|---|---|
| Phases 1–2 — prototype and instrument repair | Superseded | The historical v1 prototype and the mechanical re-adjudication layer built after it exposed analyst discretion. |
| Phase 3 — can an LLM serve as a behavioral subject? | Sealed record | Bare GPT-4.1 in repeated PD, framing, and RPS: 3 supported, 6 refuted, 1 inconclusive; X1 flips the corner. |
| Phase 4 — representation, counterfactuals, adversaries, drift | Sealed record | Registry v3, 250 sealed arms, 2,864 runs: the S2 switch, label-swap conflict, corner-confounded δ-assays, adversary suite, and the sentinel that caught endpoint drift. |
| Phase 5 — the sixteen-persona panel | Sealed record | 16 sealed persona sentences × 6 conditions × 3 temperatures; 1,712 runs; both author predictions with teeth failed. |
| Phase 6 — the preregistered replication (not yet run) | Prospective | Forbidden before publication by the scope seal; its design requirements are already written. |
The through-line
Phase 1's prototype exposed analyst discretion → Phase 2 mechanized adjudication. Phase 3 found the bare subject at corners everywhere, with a paraphrase flip (X1) showing wording dominated the incentive manipulation → Phase 4 mapped representation: localized the flip to one sentence (X2/S2), decoupled labels from payoffs (D2), found δ-assays corner-confounded (E), and caught live endpoint drift with behavioral sentinels. Phase 5 asked whether cheap persona conditioning produces an interior, incentive-sensitive population — and answered with the composition result the paper reports. Phase 6 is the preregistered replication the scope seal deferred until after publication.
Scale and integrity
| Quantity | Value |
|---|---|
| Archived completed runs (all phases) | 5,505 |
| Confirmatory replay contract | 4,916 runs (+3 legacy diagnostics = 4,919) |
| Round events / seat decisions | 54,276 / 108,552 |
| Phase 4–5 ledger | 30,530 calls · 13,141,675 input tokens · 45,247 output tokens |
| Invalid trials in Phases 4–5 | 0 (24 provider-failure partials disclosed individually) |
| Registered author predictions refuted | 12 |
| Process-failure instances publicly ledgered | 22 |