Synthetic Players

Open computational behavioral science · final release · August 2026

Passing Coarse Marginal Checks Can Be Cheap

Persona mixtures and imprecise treatment-response estimates in an LLM persona panel · Yohei Nakajima, Untapped Capital · PDF · arXiv source · repository · review record

Prefer a walkthrough? Read the whole study as a Q&A — including the findings that didn't make the paper.

LLMs are increasingly used as synthetic research participants and validated by whether their marginal responses resemble human data. A fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations met preregistered broad-reference condition-mean criteria in three of four repeated-game cells — while the treatment-response object those checks might be taken to validate stayed loosely bounded, and much of the apparent behavioral diversity was between-prompt composition of empirically corner-concentrated policies. The registered marginal criteria could be passed without precisely estimating the treatment response. One fixed model-prompt panel; no human-substitutability claim.

3 / 4repeated-game cells inside the preregistered band; sole miss 0.011 below
47–71%prior-sensitive median between-prompt share (α=1 → Jeffreys); plug-in 85–96%
+0.083 / +0.078treatment contrasts, 95% intervals ≈ [−0.18, +0.33]
0/40 → 37/40cooperation after one wording-and-position operation
4,919archived runs replayed by the public capsule — 4,916 confirmatory + 3 diagnostics

Researchers have started using AI models as stand-ins for human study participants — it's fast and nearly free. The usual quality check asks: does the AI's behavior look statistically human? This project built sixteen simple AI “personas” (one sentence each — “You are Harper, a 61-year-old landscape gardener… competitive, patient, and risk-averse”), had them play cooperation games thousands of times, and found something uncomfortable: the panel passed the standard looks-human checks while the thing experiments actually exist to measure — how behavior shifts when you change the incentives — stayed basically unmeasured. Looking human is cheap. Reacting like humans is a different, harder property, and the popular checks don't test it.

3 / 4of the game setups passed the “looks human” test — by the usual standard, the panel worked
≈ 0clear reaction when the incentive to cooperate got 9× stronger — the change was tiny and uncertain
1 sentencerewording one line flipped the base model from 0% to 92% cooperation
1 worda label reading “Defect” beat the actual money on the table
4,919archived game runs anyone can re-verify on a laptop — no AI account needed

The result in four sentences

Coarse marginal checks were passed while the registered treatment response remained imprecisely estimated — the slack in the checks is exactly where the response lives (Proposition A). The panel's apparent diversity is largely between-prompt composition of corner-concentrated policies, and whether that share is “dominant” is prior-dependent. Representation, not incentives, controlled the corners: one sentence operation moved bare cooperation from 0/40 to 37/40 while ordinary wording changes did nothing, and a displayed label overrode payoff dominance. The favored persona-level result, p13, was demoted to a replication target after external review exposed the missing family control — and the P5-2 posterior's proximity to its 0.205 boundary reading is likewise prior-dependent.

Why this is interesting

Cheap AI “participants” are already informing real decisions — pretesting surveys, products, and policy messages. This project shows the standard quality bar can be cleared by accident: a panel can look statistically human while nobody has actually checked whether it reacts to changed conditions the way the test seems to promise. It's like hiring an actor who nails the accent but ignores the script — convincing at a glance, wrong for the job.

The “diversity” was a trick of mixing. Almost every persona was locked into a habit — always cooperate or always defect. Stir eight of one and eight of the other together and the averages come out human-shaped, even though no individual is weighing the decision. That's the “cheap pass” in the title: variety that comes from the recipe, not from anyone's judgment (the decomposition).

Words beat money. Rewriting one sentence about the game continuing took the base model from never cooperating to almost always (0/40 → 37/40), while a battery of ordinary wording tweaks did nothing at all. And when a label said “Defect” on the objectively better-paying choice, the model followed the word, not the money. Behavior a sentence can rewrite isn't a stable synthetic person — which matters for safety, not just science.

The team caught its own best result being too good. One persona — Harper, the 61-year-old landscape gardener — seemed to genuinely respond to incentives. Outside reviewers showed the test behind that headline had a statistical hole, so the paper demotes its own favorite finding to “needs a proper replication.” The mistakes, the refuted predictions (twelve of them), and every correction are part of the published record — you can watch the science self-correct in the open.

Claims ledger

Every registered predicate and citable secondary, with its current evidentiary status. Each row is a page; each page links its evidence, analyses, figures, and review history. Full tier definitions on the claims page.

ClaimStatusOne-line result
P3-A3 / broad marginal checks — 3 of 4 cells in bandRegistered · verdict passRegistered broad-reference cooperation band [0.36, 0.63]; sole miss 0.011 below the lower bound.
Between-prompt composition — substantial, dominance prior-dependentPrior-sensitiveMedian between-prompt share 63–71% (Jeffreys α=.5), 47–53% (α=1); plug-in 85–96%.
Treatment response — small points, wide intervalsImprecisely estimatedContrasts +0.083 / +0.078 with conservative simultaneous 95% intervals [−0.171,+0.330] / [−0.181,+0.330].
P5-1a — corner-mixture predicateMethod-sensitiveRegistered support condition passes under 2 of 3 census methods: 3/32, 2/32, 5/32 interior.
P5-1b — between-persona dispersion checkpointRegistered · verdict passCorrected between-prompt SDs 0.418–0.478 cleared frozen thresholds 0.309/0.234; retained as permissive.
P5-2 — persona-direction vs task-text classificationRegistered · mixed recordPooled 45/352 = 0.128 task-consistent; historical verdict preserved; Bayesian proximity to 0.20 is prior-dependent.
P5-3(a) / p13 — the demoted headlineReplication target0.333→0.750 under the frozen rule, but no family control; exact-gate family is underpowered. Replication target.
P5-3(b) — dominated-option rejectionRegistered · verdict passAll 24 evaluable lanes rejected the bare configuration's dominated swap choice; minimum lower bound 0.462.
X2 / S2 switch — one sentence operation, 0/40 → 37/40Registered · verdict passA single wording-and-position operation on the continuation sentence moved bare cooperation across the range.
D1 — ordinary one-shot wording main effect (null)Registered · not supported+0.0063 (SE 0.0210; Holm-adjusted p=1.00) across the 640-episode one-shot battery.
Label-swap conflict — semantic label overrides payoff dominanceRegistered · verdict passBare GPT-4.1 chose the cooperation-worded option 0/40, taking the payoff-dominated role when it carried “Defect.”
Leaning-stratum gaps — 0.51 to 0.72DescriptivePreregistered trait-rule strata differ by 0.510–0.719 across conditions; bundle contrasts, not trait causality.
RPS adversary suite — exploitability is opponent-contingentRegistered · mixed recordn-gram2 earned +0.215/round vs GPT-4.1 (Holm-surviving); the WSLS-targeter designed from its signature earned +0.008.
RPS role-attached asymmetry — registered direction reversedRegistered · not supportedFirst-minus-rock contrast −0.181 (prediction was positive); Gemini mirror +0.243. Vendor-specific asymmetry.
Cross-vendor label–payoff dissociationDescriptiveGPT-4.1 followed the token DEFECT even when payoff-dominated; Gemini mostly followed payoff dominance.
Temperature secondary — matched-lattice entropy declineDescriptivePooled entropy 0.831 → 0.782 → 0.770 bits at T=0.7/1.0/1.3 on the identical 13-unit lattice.
Endpoint drift and subject eligibilityProcedural recordGemini sentinel fell from 10/10 to 6/10–7/10 on an unversioned endpoint; Claude Haiku failed the entry gate.

The findings, in plain terms

  • Passed the resemblance test: in 3 of 4 game setups, group averages landed inside the pre-registered “human range” — the checks the field actually uses.
  • Barely reacted to incentives: making future rounds 9× more likely moved cooperation by about 8 points out of 100, with uncertainty so wide the true effect could plausibly be negative — the imprecise response.
  • Habits, not decisions: ten of sixteen personas never varied their choice in any setup; the human-looking spread came from mixing stubborn habits — the mixture result.
  • Wording is the control knob: one rewritten sentence flipped behavior end to end; a one-word label outweighed the payoffs — the switch and the label conflict.
  • The star witness got demoted: the single persona that seemed incentive-sensitive rests on a test with no multiple-comparisons control; it's now officially a “replication target,” not a finding.

Want the full ledger with statistical detail? Flip to Technical, or open the claims page.

Figures

The paper's five figures, each with sources and provenance.

The paper's five figures — click any of them for a guided explanation of what you're seeing.

The program

Five phases, sealed registries, mechanical adjudication, 12 refuted author predictions, 14 review rounds, 15 manuscript versions. Full timeline · phases · reviews · versions.

PhaseWhat it establishedScale
Phases 1–2prototype → mechanical re-adjudication after it exposed analyst discretion40 + 400 experiments
Phase 3bare GPT-4.1 sits at corners; one paraphrase flips 0.000 → 1.000320 runs, 5,820 calls
Phase 4the S2 switch, label/payoff conflict, corner-confounded δ-assays, sentinel catches endpoint drift2,864 runs, 20,102 calls
Phase 5the sixteen-persona panel; both author predictions with teeth failed1,712 runs, 10,428 calls
Phase 6preregistered replication — power-planned, not yet runprospective

How the study was run

Five stages over five weeks, with predictions locked in — publicly and tamper-evidently — before the data existed, and a computer (not the author) deciding pass or fail.

StagePlain-language version
Phases 1–2built the lab — and learned that human judgment sneaks into scoring, so scoring was handed to sealed, automatic rules
Phase 3the plain AI (no persona) turned out to be an extreme case: it essentially never cooperates — until a rewording flips it completely
Phase 4hunted down which words matter: found the one controlling sentence, showed labels can beat money, and a tripwire caught the AI vendor's model changing mid-study
Phase 5gave the AI sixteen one-sentence personalities and ran the full experiment — the author's two boldest predictions both failed, and that failure is the finding
Phase 6the honest follow-up: a bigger, properly powered redo, designed before any new data is collected

How this differs from prior work

The broad “realism ≠ effect accuracy” thesis is occupied; the contribution here is narrower and mechanism-level. Full map with per-paper differentiation on the related-work page.

Li & Ji establish the realism/effect divergence at survey scale; this project identifies one concrete strategic-interaction pattern behind cheap passes — how we differ. Persson et al. formalize LLM causal surrogacy; this is a registered design-side example of what coarse checks leave unidentified — details. Lin et al.'s latent-user drift can coexist with the composition pattern measured here — details. Pal et al. run nearly the same manipulations without the persona panel or audit architecture — details. Ashokkumar et al. is the strong positive counterexample, on a different estimand — details.

Hasn't someone shown this already?

Partly — and the paper says so. Large studies have shown AI simulations can look right while getting cause-and-effect wrong, and one impressive study shows AI can forecast experiment outcomes surprisingly well. What's new here is the mechanism, caught in the act: a specific, common recipe (one-sentence personas) that passes the checks by mixing stuck habits — plus receipts. Predictions were locked before data, every AI call was archived, and the whole record replays on your machine. The related-work page compares this paper to each neighboring study in a sentence or two.

Reproduce the confirmatory record

Don't take our word for it

One command replays every confirmatory run — 4,916 confirmatory plus three legacy diagnostics — byte-exact from the archived, checksummed, externally timestamped databases. No credentials, no live model calls. The capsule page documents what replay does and does not prove.

Expected: CAPSULE VERIFICATION PASS — 4,919 archived Phase 3-5 runs verified.

Every archived game — all 4,919 — re-checks byte-for-byte from the published record with the three lines on the right. No AI accounts or API keys, because nothing is re-generated: your machine verifies the archive against itself, including every prompt, every response, and every scored outcome. What that does and doesn't prove.

git clone https://github.com/yoheinakajima/synthetic-players
cd synthetic-players/capsule
bash verify.sh

Materials

ArtifactWhat it is
Paper PDF + arXiv packagecanonical 19-page PDF, byte-identical Markdown, minimal PDFLaTeX zip — all hash-pinned
Replay capsuleone-command zero-credential verification of 4,919 archived runs
Checksums, timestamps, sealsSHA-256 manifests, OpenTimestamps Bitcoin anchors, sealed tags and releases
Persona tableall sixteen sealed sentences with construction rule and hashes
Dead-predictions ledger12 refuted author predictions, adjudicated by sealed code
Review archiveevery external critique, disposition matrix, and role disclosure