Positioning
Related work — occupied territory and precise differentiation
What is already established, what collides, and the narrow triangle this paper defends.
The defensible novelty triangle
1. A registered strategic-interaction example where a fixed persona panel passes coarse marginal checks while continuation-probability estimates stay small and imprecise. 2. The mechanism-level pattern: dispersion carried largely by between-prompt composition of empirically corner-concentrated policies. 3. The credibility layer: registration, provenance, replay, mechanical adjudication, and public post-adjudication correction. Explicitly not claimed: first demonstration of realism/effect divergence, drift-free panels, human interiority, trait causality, or a p13 capability finding.
Closest occupied territory
| Work | What it establishes | |
|---|---|---|
| Li & Ji 2026 | When simulations look right but causal effects go wrong: LLMs as behavioral simulators | source ↗ |
| Ashokkumar et al. 2026 | Large language models can predict the results of social science experiments | source ↗ |
| Persson et al. 2026 | Statistical foundations of LLM-based A/B testing: a surrogacy framework | source ↗ |
| Lin et al. 2026 | The illusion of intervention: your LLM-simulated experiment is an observational study | source ↗ |
| Xie et al. 2026 | Evaluating the statistical realism of LLM-generated social science data (SSDataBench) | source ↗ |
| Harry et al. 2026 | Beyond fixed psychological personas: state beats trait, but language models are state-blind | source ↗ |
| Xiao et al. 2026 | The chameleon's limit: persona collapse and homogenization in LLMs | source ↗ |
Direct strategic-behavior collisions
| Work | Collision | |
|---|---|---|
| Akata et al. 2025 | Playing repeated games with large language models | source ↗ |
| Pal et al. 2026 | Strategies of cooperation and defection in five large language models | source ↗ |
| Georgousis et al. 2026 | Evaluating counterfactual strategic reasoning in large language models | source ↗ |
| Mousavi Davoudi et al. 2026 | Same game, different story: a strategic-robustness benchmark | source ↗ |
| Mei et al. 2024 | A Turing test of whether AI chatbots are behaviorally similar to humans | source ↗ |
Synthetic participants and personas
| Work | Relation | |
|---|---|---|
| Bisbee et al. 2024 | Synthetic replacements for human survey data? The perils of large language models | source ↗ |
| Boelaert et al. 2025 | Machine bias: how do generative language models answer opinion polls? | source ↗ |
| Anthis et al. 2025 | Position: LLM social simulations are a promising research method | source ↗ |
| Hullman et al. 2026 | This human study did not involve human subjects: validating LLM simulations | source ↗ |
| Park et al. 2024 | LLM agents grounded in self-reports enable general-purpose simulation of individuals | source ↗ |
| Argyle et al. 2023 | Out of one, many: using language models to simulate human samples | source ↗ |
| Horton 2023 | Large language models as simulated economic agents (Homo Silicus) | source ↗ |
| Batzner et al. 2025 | Whose personae? Synthetic persona experiments and pathways to transparency | source ↗ |
| Sclar et al. 2024 | Quantifying language models' sensitivity to spurious features in prompt design | source ↗ |
| Shanahan et al. 2023 | Role play with large language models | source ↗ |
Comparators and classical lineages
| Work | Use here | |
|---|---|---|
| Dal Bó & Fréchette 2011 | The evolution of cooperation in infinitely repeated games | source ↗ |
| Lucas 1976 | Econometric policy evaluation: a critique | source ↗ |
| Cronbach & Meehl 1955 | Construct validity in psychological tests | source ↗ |
| ICH E10 / Temple & Ellenberg 2000 | Assay sensitivity in controlled trials | source ↗ |
| Windrum et al. 2007 / Grimm et al. 2005 | Agent-based model validation and equifinality | source ↗ |
| Statistical methods | Statistical foundations used by the analyses | source ↗ |
Sources: the paper's §2 and References, plus the archived literature map and novelty-relationships documents (working research maps, not verdict-bearing).