Open computational behavioral science · final release · August 2026
Passing Coarse Marginal Checks Can Be Cheap
Prefer a walkthrough? Read the whole study as a Q&A — including the findings that didn't make the paper.
LLMs are increasingly used as synthetic research participants and validated by whether their marginal responses resemble human data. A fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations met preregistered broad-reference condition-mean criteria in three of four repeated-game cells — while the treatment-response object those checks might be taken to validate stayed loosely bounded, and much of the apparent behavioral diversity was between-prompt composition of empirically corner-concentrated policies. The registered marginal criteria could be passed without precisely estimating the treatment response. One fixed model-prompt panel; no human-substitutability claim.
Researchers have started using AI models as stand-ins for human study participants — it's fast and nearly free. The usual quality check asks: does the AI's behavior look statistically human? This project built sixteen simple AI “personas” (one sentence each — “You are Harper, a 61-year-old landscape gardener… competitive, patient, and risk-averse”), had them play cooperation games thousands of times, and found something uncomfortable: the panel passed the standard looks-human checks while the thing experiments actually exist to measure — how behavior shifts when you change the incentives — stayed basically unmeasured. Looking human is cheap. Reacting like humans is a different, harder property, and the popular checks don't test it.
The result in four sentences
Coarse marginal checks were passed while the registered treatment response remained imprecisely estimated — the slack in the checks is exactly where the response lives (Proposition A). The panel's apparent diversity is largely between-prompt composition of corner-concentrated policies, and whether that share is “dominant” is prior-dependent. Representation, not incentives, controlled the corners: one sentence operation moved bare cooperation from 0/40 to 37/40 while ordinary wording changes did nothing, and a displayed label overrode payoff dominance. The favored persona-level result, p13, was demoted to a replication target after external review exposed the missing family control — and the P5-2 posterior's proximity to its 0.205 boundary reading is likewise prior-dependent.
Why this is interesting
Cheap AI “participants” are already informing real decisions — pretesting surveys, products, and policy messages. This project shows the standard quality bar can be cleared by accident: a panel can look statistically human while nobody has actually checked whether it reacts to changed conditions the way the test seems to promise. It's like hiring an actor who nails the accent but ignores the script — convincing at a glance, wrong for the job.
The “diversity” was a trick of mixing. Almost every persona was locked into a habit — always cooperate or always defect. Stir eight of one and eight of the other together and the averages come out human-shaped, even though no individual is weighing the decision. That's the “cheap pass” in the title: variety that comes from the recipe, not from anyone's judgment (the decomposition).
Words beat money. Rewriting one sentence about the game continuing took the base model from never cooperating to almost always (0/40 → 37/40), while a battery of ordinary wording tweaks did nothing at all. And when a label said “Defect” on the objectively better-paying choice, the model followed the word, not the money. Behavior a sentence can rewrite isn't a stable synthetic person — which matters for safety, not just science.
The team caught its own best result being too good. One persona — Harper, the 61-year-old landscape gardener — seemed to genuinely respond to incentives. Outside reviewers showed the test behind that headline had a statistical hole, so the paper demotes its own favorite finding to “needs a proper replication.” The mistakes, the refuted predictions (twelve of them), and every correction are part of the published record — you can watch the science self-correct in the open.
Claims ledger
Every registered predicate and citable secondary, with its current evidentiary status. Each row is a page; each page links its evidence, analyses, figures, and review history. Full tier definitions on the claims page.
| Claim | Status | One-line result |
|---|---|---|
| P3-A3 / broad marginal checks — 3 of 4 cells in band | Registered · verdict pass | Registered broad-reference cooperation band [0.36, 0.63]; sole miss 0.011 below the lower bound. |
| Between-prompt composition — substantial, dominance prior-dependent | Prior-sensitive | Median between-prompt share 63–71% (Jeffreys α=.5), 47–53% (α=1); plug-in 85–96%. |
| Treatment response — small points, wide intervals | Imprecisely estimated | Contrasts +0.083 / +0.078 with conservative simultaneous 95% intervals [−0.171,+0.330] / [−0.181,+0.330]. |
| P5-1a — corner-mixture predicate | Method-sensitive | Registered support condition passes under 2 of 3 census methods: 3/32, 2/32, 5/32 interior. |
| P5-1b — between-persona dispersion checkpoint | Registered · verdict pass | Corrected between-prompt SDs 0.418–0.478 cleared frozen thresholds 0.309/0.234; retained as permissive. |
| P5-2 — persona-direction vs task-text classification | Registered · mixed record | Pooled 45/352 = 0.128 task-consistent; historical verdict preserved; Bayesian proximity to 0.20 is prior-dependent. |
| P5-3(a) / p13 — the demoted headline | Replication target | 0.333→0.750 under the frozen rule, but no family control; exact-gate family is underpowered. Replication target. |
| P5-3(b) — dominated-option rejection | Registered · verdict pass | All 24 evaluable lanes rejected the bare configuration's dominated swap choice; minimum lower bound 0.462. |
| X2 / S2 switch — one sentence operation, 0/40 → 37/40 | Registered · verdict pass | A single wording-and-position operation on the continuation sentence moved bare cooperation across the range. |
| D1 — ordinary one-shot wording main effect (null) | Registered · not supported | +0.0063 (SE 0.0210; Holm-adjusted p=1.00) across the 640-episode one-shot battery. |
| Label-swap conflict — semantic label overrides payoff dominance | Registered · verdict pass | Bare GPT-4.1 chose the cooperation-worded option 0/40, taking the payoff-dominated role when it carried “Defect.” |
| Leaning-stratum gaps — 0.51 to 0.72 | Descriptive | Preregistered trait-rule strata differ by 0.510–0.719 across conditions; bundle contrasts, not trait causality. |
| RPS adversary suite — exploitability is opponent-contingent | Registered · mixed record | n-gram2 earned +0.215/round vs GPT-4.1 (Holm-surviving); the WSLS-targeter designed from its signature earned +0.008. |
| RPS role-attached asymmetry — registered direction reversed | Registered · not supported | First-minus-rock contrast −0.181 (prediction was positive); Gemini mirror +0.243. Vendor-specific asymmetry. |
| Cross-vendor label–payoff dissociation | Descriptive | GPT-4.1 followed the token DEFECT even when payoff-dominated; Gemini mostly followed payoff dominance. |
| Temperature secondary — matched-lattice entropy decline | Descriptive | Pooled entropy 0.831 → 0.782 → 0.770 bits at T=0.7/1.0/1.3 on the identical 13-unit lattice. |
| Endpoint drift and subject eligibility | Procedural record | Gemini sentinel fell from 10/10 to 6/10–7/10 on an unversioned endpoint; Claude Haiku failed the entry gate. |
The findings, in plain terms
- Passed the resemblance test: in 3 of 4 game setups, group averages landed inside the pre-registered “human range” — the checks the field actually uses.
- Barely reacted to incentives: making future rounds 9× more likely moved cooperation by about 8 points out of 100, with uncertainty so wide the true effect could plausibly be negative — the imprecise response.
- Habits, not decisions: ten of sixteen personas never varied their choice in any setup; the human-looking spread came from mixing stubborn habits — the mixture result.
- Wording is the control knob: one rewritten sentence flipped behavior end to end; a one-word label outweighed the payoffs — the switch and the label conflict.
- The star witness got demoted: the single persona that seemed incentive-sensitive rests on a test with no multiple-comparisons control; it's now officially a “replication target,” not a finding.
Want the full ledger with statistical detail? Flip to Technical, or open the claims page.
Figures
The paper's five figures, each with sources and provenance.
The paper's five figures — click any of them for a guided explanation of what you're seeing.
The program
Five phases, sealed registries, mechanical adjudication, 12 refuted author predictions, 14 review rounds, 15 manuscript versions. Full timeline · phases · reviews · versions.
| Phase | What it established | Scale |
|---|---|---|
| Phases 1–2 | prototype → mechanical re-adjudication after it exposed analyst discretion | 40 + 400 experiments |
| Phase 3 | bare GPT-4.1 sits at corners; one paraphrase flips 0.000 → 1.000 | 320 runs, 5,820 calls |
| Phase 4 | the S2 switch, label/payoff conflict, corner-confounded δ-assays, sentinel catches endpoint drift | 2,864 runs, 20,102 calls |
| Phase 5 | the sixteen-persona panel; both author predictions with teeth failed | 1,712 runs, 10,428 calls |
| Phase 6 | preregistered replication — power-planned, not yet run | prospective |
How the study was run
Five stages over five weeks, with predictions locked in — publicly and tamper-evidently — before the data existed, and a computer (not the author) deciding pass or fail.
| Stage | Plain-language version |
|---|---|
| Phases 1–2 | built the lab — and learned that human judgment sneaks into scoring, so scoring was handed to sealed, automatic rules |
| Phase 3 | the plain AI (no persona) turned out to be an extreme case: it essentially never cooperates — until a rewording flips it completely |
| Phase 4 | hunted down which words matter: found the one controlling sentence, showed labels can beat money, and a tripwire caught the AI vendor's model changing mid-study |
| Phase 5 | gave the AI sixteen one-sentence personalities and ran the full experiment — the author's two boldest predictions both failed, and that failure is the finding |
| Phase 6 | the honest follow-up: a bigger, properly powered redo, designed before any new data is collected |
How this differs from prior work
The broad “realism ≠ effect accuracy” thesis is occupied; the contribution here is narrower and mechanism-level. Full map with per-paper differentiation on the related-work page.
Li & Ji establish the realism/effect divergence at survey scale; this project identifies one concrete strategic-interaction pattern behind cheap passes — how we differ. Persson et al. formalize LLM causal surrogacy; this is a registered design-side example of what coarse checks leave unidentified — details. Lin et al.'s latent-user drift can coexist with the composition pattern measured here — details. Pal et al. run nearly the same manipulations without the persona panel or audit architecture — details. Ashokkumar et al. is the strong positive counterexample, on a different estimand — details.
Hasn't someone shown this already?
Partly — and the paper says so. Large studies have shown AI simulations can look right while getting cause-and-effect wrong, and one impressive study shows AI can forecast experiment outcomes surprisingly well. What's new here is the mechanism, caught in the act: a specific, common recipe (one-sentence personas) that passes the checks by mixing stuck habits — plus receipts. Predictions were locked before data, every AI call was archived, and the whole record replays on your machine. The related-work page compares this paper to each neighboring study in a sentence or two.
Reproduce the confirmatory record
Don't take our word for it
One command replays every confirmatory run — 4,916 confirmatory plus three legacy diagnostics — byte-exact from the archived, checksummed, externally timestamped databases. No credentials, no live model calls. The capsule page documents what replay does and does not prove.
Expected: CAPSULE VERIFICATION PASS — 4,919 archived Phase 3-5 runs verified.
Every archived game — all 4,919 — re-checks byte-for-byte from the published record with the three lines on the right. No AI accounts or API keys, because nothing is re-generated: your machine verifies the archive against itself, including every prompt, every response, and every scored outcome. What that does and doesn't prove.
git clone https://github.com/yoheinakajima/synthetic-players
cd synthetic-players/capsule
bash verify.shMaterials
| Artifact | What it is |
|---|---|
| Paper PDF + arXiv package | canonical 19-page PDF, byte-identical Markdown, minimal PDFLaTeX zip — all hash-pinned |
| Replay capsule | one-command zero-credential verification of 4,919 archived runs |
| Checksums, timestamps, seals | SHA-256 manifests, OpenTimestamps Bitcoin anchors, sealed tags and releases |
| Persona table | all sixteen sealed sentences with construction rule and hashes |
| Dead-predictions ledger | 12 refuted author predictions, adjudicated by sealed code |
| Review archive | every external critique, disposition matrix, and role disclosure |




