Synthetic Players

Ledger

Claims and their evidentiary tiers

Every registered predicate, its frozen verdict, and its current scientific status — kept deliberately distinct.

The program's rule: historical mechanical verdicts stay visible verbatim; later analyses are additive and explicitly labeled. Twelve registered author predictions were refuted and published (the dead-predictions ledger); the p13 demotion is an inferential correction, not a refutation, and is deliberately not counted among them.

Tier vocabulary

TierMeaning
Registered · verdict pass/failfrozen mechanical verdicts, adjudicated by sealed code before interpretation; historical labels never rewritten
Method-/prior-sensitiveclassification changes across defensible interval constructions or symmetric priors; continuous estimates primary
Replication targetneither prospectively confirmed nor decisively disconfirmed; next test must be prospectively powered
Descriptivereported without inferential weight (cross-vendor tier, non-registered comparisons)
Post-adjudicationcomputed after the seal in response to review; cannot create prospective confirmation
Proceduralmonitoring and process records, not behavioral findings

All claims

ClaimStatusOne-line result
P3-A3 / broad marginal checks — 3 of 4 cells in bandRegistered · verdict passRegistered broad-reference cooperation band [0.36, 0.63]; sole miss 0.011 below the lower bound.
Between-prompt composition — substantial, dominance prior-dependentPrior-sensitiveMedian between-prompt share 63–71% (Jeffreys α=.5), 47–53% (α=1); plug-in 85–96%.
Treatment response — small points, wide intervalsImprecisely estimatedContrasts +0.083 / +0.078 with conservative simultaneous 95% intervals [−0.171,+0.330] / [−0.181,+0.330].
P5-1a — corner-mixture predicateMethod-sensitiveRegistered support condition passes under 2 of 3 census methods: 3/32, 2/32, 5/32 interior.
P5-1b — between-persona dispersion checkpointRegistered · verdict passCorrected between-prompt SDs 0.418–0.478 cleared frozen thresholds 0.309/0.234; retained as permissive.
P5-2 — persona-direction vs task-text classificationRegistered · mixed recordPooled 45/352 = 0.128 task-consistent; historical verdict preserved; Bayesian proximity to 0.20 is prior-dependent.
P5-3(a) / p13 — the demoted headlineReplication target0.333→0.750 under the frozen rule, but no family control; exact-gate family is underpowered. Replication target.
P5-3(b) — dominated-option rejectionRegistered · verdict passAll 24 evaluable lanes rejected the bare configuration's dominated swap choice; minimum lower bound 0.462.
X2 / S2 switch — one sentence operation, 0/40 → 37/40Registered · verdict passA single wording-and-position operation on the continuation sentence moved bare cooperation across the range.
D1 — ordinary one-shot wording main effect (null)Registered · not supported+0.0063 (SE 0.0210; Holm-adjusted p=1.00) across the 640-episode one-shot battery.
Label-swap conflict — semantic label overrides payoff dominanceRegistered · verdict passBare GPT-4.1 chose the cooperation-worded option 0/40, taking the payoff-dominated role when it carried “Defect.”
Leaning-stratum gaps — 0.51 to 0.72DescriptivePreregistered trait-rule strata differ by 0.510–0.719 across conditions; bundle contrasts, not trait causality.
RPS adversary suite — exploitability is opponent-contingentRegistered · mixed recordn-gram2 earned +0.215/round vs GPT-4.1 (Holm-surviving); the WSLS-targeter designed from its signature earned +0.008.
RPS role-attached asymmetry — registered direction reversedRegistered · not supportedFirst-minus-rock contrast −0.181 (prediction was positive); Gemini mirror +0.243. Vendor-specific asymmetry.
Cross-vendor label–payoff dissociationDescriptiveGPT-4.1 followed the token DEFECT even when payoff-dominated; Gemini mostly followed payoff dominance.
Temperature secondary — matched-lattice entropy declineDescriptivePooled entropy 0.831 → 0.782 → 0.770 bits at T=0.7/1.0/1.3 on the identical 13-unit lattice.
Endpoint drift and subject eligibilityProcedural recordGemini sentinel fell from 10/10 to 6/10–7/10 on an unversioned endpoint; Claude Haiku failed the entry gate.

Dependency spine

The paper's argument rests on four load-bearing links: the marginal checks license nothing about response beyond band arithmetic (Proposition A); the composition claim depends on which uncertainty view conditions on boundary concentration; the p13 story rests on the family audits and the attainability floor; and the P5-2 classification is carried by the word/payoff-confounded swap cell. The archived claim-dependency audit maps every quotation of these results to its correction rule.