Synthetic Players

PhaseSealed record

Phase 3 — can an LLM serve as a behavioral subject?

Bare GPT-4.1 in repeated PD, framing, and RPS: 3 supported, 6 refuted, 1 inconclusive; X1 flips the corner.

Complete 2026-07-24 · registry phase3-v1/v2 · 320 LLM runs + 20 baselines · claims registered before data

Phase 3 reportPreregistrationLayer-2 statistics

Bare GPT-4.1 (temperature 0.7, max_tokens 16, no persona) across three families: A — random-termination repeated PD at δ ∈ {.10, .50, .75, .90} with a payoff isomorph; B — one-shot framing (Community / Wall Street / neutral); C — 50-round RPS versus a pattern tracker, Nash mixing, and self-play. 320 LLM runs, 5,820 calls, plus 20 deterministic baselines.

Registered verdicts

PredicateVerdictKey number
P3-A1 shadow of the futurerefuted0.000 cooperation at both δ
P3-A2 risk-dominance separationrefuted0.000 vs 0.000
P3-A3 human band [0.36, 0.63]refuted0.000
P3-A4 isomorph invariancerefutedfails the separation limb
P3-B1 framing directionsupported0.175 vs 0.000, CI [0.061, 0.290]
P3-B2 framing magnitudeinconclusiveedge rule at Wall Street = 0
P3-B3 neutral interiorsupported0.000 ≤ 0.000 ≤ 0.175
P3-C1 RPS rock bandrefutedrock 0.80 ∉ [0.33, 0.40]
P3-C2 win-stay / lose-shiftsupportedP(shift|lose) 0.974
P3-C3 tracker exploits LLMrefuted, sign reversed−0.103; subject beat the tracker
P3-X1 paraphrase robustnessrefuted0.000 → 1.000 under two rewordings, same seeds

The approved headline: prompt wording dominated the tested incentive manipulation and rendered single-wording behavioral inference non-identifiable. X1’s total corner flip is what motivated Phase 4’s representation program. All 320 runs replay bit-exact with zero live calls; the disclosed Phase 3 gap (no provider response IDs) is stated in the paper.