Can LLM agents stand in for people in social simulations?
Not reliably on this evidence. Across eight models from three families, simulated conversations were more repetitive, more positive and less representative of real occupations than human dialogue records. The peer-reviewed benchmark identifies a directional bias—not a verdict on every model, population or simulation task.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The researchers compared eight GPT, Llama and DeepSeek models with human conversations from TopicalChat and PersonaChat across role distribution, semantic similarity, keyword persistence, sentiment and linguistic patterns.
- 2The sentiment analysis covered 178,824 LLM utterances and 319,816 human utterances; the model conversations were significantly more positive overall, while role generation overrepresented professional and managerial occupations.
- 3The benchmark uses structured, cyclic chatrooms and existing crowd-worker dialogue datasets. It does not show that every LLM simulation fails, or that the human datasets represent every society and social setting.
Research topic
Whether multi-agent LLM conversations reproduce the diversity, disagreement and role distribution found in human social interaction

The direct answer: these agents were too agreeable to be human substitutes
The eight tested large language models did not behave like neutral samples of people. In structured chatroom simulations, their conversations were more semantically repetitive, more positive and more concentrated on the opening topic than the human dialogue records used for comparison. When asked to generate social roles, the models also filled their synthetic worlds disproportionately with professionals and managers while underrepresenting elementary and agricultural work.
That pattern matters because researchers and organisations increasingly use LLM agents to explore opinions, group dynamics and possible responses to policy or products. A simulation that smooths away disagreement can make a proposal look more consensual than it is. A synthetic population that favours prestigious jobs can conceal the people most exposed to a decision. The paper's practical warning is therefore specific: do not treat outputs from aligned chat models as observations of society without validating them against relevant humans.
The result is not a general ban on agent-based research. It is evidence that a common implementation—role-conditioned language models taking turns in a chat—can reproduce the socially desirable behaviour rewarded during model training rather than the variation found in human interaction. Models can still help generate hypotheses or stress-test mechanisms, provided their behaviour is measured as model behaviour and not relabelled as public opinion.[1][2]
What the researchers tested
The study covered eight models from three families: GPT-4o, o3 and GPT-3.5 Turbo Instruct; Llama 3.1 8B, Llama 3.3 70B and Llama 4 17B-16E; and DeepSeek Chat and DeepSeek Reasoner. GPT and DeepSeek systems were accessed through APIs, while the Llama models ran locally on two RTX 4090 GPUs. The researchers used the models' default temperature and sampling settings, an important choice because different settings could change diversity and repetition.
Agents were assigned social roles and placed in turn-taking conversations with between two and eight participants for the principal dialogue comparisons. Topics came from three sources: model-generated prompts, keyword combinations drawn from TopicalChat and debate topics used in earlier research. The team then assessed five dimensions—role distribution, inter-agent and intra-agent semantic similarity, keyword persistence, emotional tone and broader linguistic patterns.
Human baselines came from TopicalChat and PersonaChat, both crowd-sourced English-language dialogue collections. The TopicalChat training set contained 188,378 utterances across 8,628 dialogues, with conversations required to last at least 20 turns. Across the paper's sentiment comparison, the reported denominator was 178,824 LLM utterances and 319,816 human utterances. These are large text samples, but they are not a probability sample of human societies.[1][2]
Three distortions produced the 'utopian' pattern
First came role bias. The researchers mapped generated occupations to the International Standard Classification of Occupations and compared them with International Labour Organization estimates. Llama 3.1 generated 5,300 roles and each other model generated 1,100. Professionals accounted for 26.6% to 62.2% of generated roles, compared with 10.4% of the global workforce estimate used by the authors. Elementary and agricultural, forestry and fishery work represented 40.2% of that benchmark but only 0.7% to 6.4% of generated roles.
Second came repetition and primacy. BERT-based cosine similarity was higher between consecutive model utterances than in the human dialogue datasets, both across agents and within one agent's turns. Keywords appearing early in a model conversation also retained more influence. The effect can create a coherent exchange, but coherence is not the same as authentic group reasoning: people interrupt, reinterpret, abandon premises and introduce competing concerns.
Third came positivity. Using VADER sentiment scores, the authors found model utterances significantly more positive overall than the human records. The paper reports p values below 0.0001 for the overall comparison, although o3 on TopicalChat was an exception and did not differ significantly. Linguistic analysis with LIWC also indicated fewer negations and less of the emotional variation present in the human data. The claim is about measured text patterns, not whether politeness itself is undesirable.[1][2]
Why the human comparison is useful but not universal
TopicalChat and PersonaChat give the study a real human baseline rather than comparing one model with another. That is a strength. Yet both datasets were produced through designed crowd-work tasks, mostly as orderly written exchanges between strangers following instructions. They do not capture a family argument, a union meeting, a multilingual neighbourhood, non-verbal communication or the power dynamics of a workplace decision.
The authors acknowledge that their chatrooms were simplified, tension-free and non-goal-oriented. Cyclic turn-taking prevents interruptions and selective participation. Agents did not share histories, physical settings or scarce resources, and the design omitted subgroup formation. Those features are often exactly where conflict, coalition-building and unequal influence emerge.
The evaluation tools introduce another layer of abstraction. Cosine similarity, KeyBERT, VADER and LIWC provide repeatable quantities, but none can determine whether a conversation is socially authentic in full. The proposed link between preference alignment and socially desirable behaviour is correlational: the study analysed preference datasets and model outputs, but it did not experimentally isolate a single training intervention as the cause.[1][2]
What this changes for research, policy and product testing
A team using synthetic agents should publish a validation table before publishing a simulated public response. That table should compare the agents with a relevant human sample on the variables that matter: demographic and occupational composition, disagreement, topic movement, attrition, minority positions and sensitivity to the prompt. A model that matches one dataset on average can still fail for subgroups or under a different decision structure.
For policy work, simulated consensus should never substitute for consultation with affected people. The paper shows how aligned models can make a discussion appear harmonious and professionally skewed. That is especially risky where a decision affects workers, minoritised communities or people whose language is poorly represented in training data. LLM simulations may help officials discover questions to ask; they cannot supply democratic legitimacy or lived experience.
Product researchers face a similar boundary. Agents can cheaply expose obvious wording problems and generate edge-case hypotheses, but willingness to buy, disclose data or tolerate harm must be measured with real users. If an agent study is used at all, the model names, versions, prompts, sampling settings, run dates and repeated-run variability should travel with the result so readers can distinguish a reproducible experiment from a persuasive demonstration.[1][2]
Funding, provenance and what remains unproven
The peer-reviewed article was published on 7 October 2026 after an earlier preprint appeared on 24 October 2025. The journal page says the work was supported by the China Postdoctoral Science Foundation under grant 2025M782534, and the authors declare no competing interests. The research institutions are the China University of Mining and Technology and the Institute of Software at the Chinese Academy of Sciences.
No new human participants were recruited by the authors; the human comparison used existing dialogue datasets. This avoids presenting experimental risk to newly enrolled people, but it also means the study cannot ask participants why they disagreed, how authentic they found the task or whether dataset workers match the population a later simulation intends to represent.
The findings do not establish that reasoning models are reliable human proxies, even though o3 and DeepSeek Reasoner sometimes produced less repetitive or less positive text. Nor do they show that every newer model or carefully calibrated agent architecture will behave the same way. Eight models across three major families provide breadth, not comprehensive coverage of model versions, languages and social-simulation designs.[1][2]
What evidence would change the assessment
Confidence would rise if independent teams preregistered simulations and compared them prospectively with human groups facing the same incentives, information and decision rules. The most informative studies would span languages and regions, include tension and resource constraints, measure interruptions and coalition formation, and report failures as carefully as average similarity scores.
Ablation experiments could test causes by varying preference alignment, system prompts, personas, memory, sampling temperature and turn-taking while holding other conditions constant. Repeated runs are essential because a single dialogue cannot reveal output variability. Researchers should also test whether calibration transfers: an intervention that improves realism on PersonaChat may not work in healthcare consent, disaster response or labour negotiations.
Until that evidence exists, the safest interpretation is narrow but consequential. These models generated simulations that were systematically cleaner, happier and more prestigious than the human records chosen for comparison. LLM agents can be objects of study and tools for exploring mechanisms; on this evidence, they should not quietly become stand-ins for people.[1][2]
What this means for people
- People affected by policy or product decisions could be misrepresented if agreeable model agents are treated as public opinion.
- Workers and lower-status occupations risk being undercounted in synthetic populations built from automatically generated roles.
- Researchers gain a concrete set of validation checks, but still need relevant human evidence before drawing social conclusions.
Global context
The study was led by institutions in China and tested major US, Chinese and open-weight model families against English-language human dialogue datasets. Its warning is globally relevant because synthetic-agent research crosses borders, but the benchmark is not a multicultural population study. Validation must be repeated in the language, institutional setting and affected population of each intended use.
What the evidence does not yet show
- The eight models came from GPT, Llama and DeepSeek families; other architectures, versions and languages may behave differently.
- The structured, cyclic chatrooms omitted interruptions, non-verbal cues, shared histories, physical context, resource constraints and goal conflict.
- TopicalChat and PersonaChat are designed crowd-worker dialogue datasets, not representative samples of all human societies or interactions.
- Automated metrics capture selected textual properties but cannot fully establish social authenticity.
- The proposed connection between training or preference alignment and the observed biases is correlational rather than causal.
What to watch next
- Prospective, preregistered comparisons in which matched human groups and LLM agents face the same decision task.
- Replication across languages, cultures, model families and high-conflict or resource-constrained settings.
- Ablations that isolate the effects of alignment, prompts, personas, memory and sampling settings.
- Standards requiring synthetic-population studies to report model versions, prompts, run dates, repeated-run variance and human validation.
- Evidence that calibration on one dialogue dataset transfers to a real policy, market or social-science use case.
Living evidence record
Impact record IAI-1FJU9GF
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
7 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 7 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
AI Risks & Safety
Can a few poisoned documents reframe medical AI answers?
Yes, in a peer-reviewed laboratory study across three biomedical corpora and five open-weight language models. Injecting five crafted documents per question often put at least three in the top five retrieved results, and one or two retrieved poisons could be enough in successful attacks.
6 min · 1 source
AI Risks & Safety
Do AI explanations prevent over-reliance?
Not in this unreviewed experiment. Explanations reduced raw acceptance but did not improve reliance calibration among novices doing a clinical-text annotation task; over-reliance tended to increase across the session.
7 min · 1 source
AI Risks & Safety
Can AI agents learn what is appropriate?
Perhaps, but the new Google-led report is a research agenda, not a tested safeguard. More than 50 contributors propose contextual policy engines, dynamic sandboxes and multi-agent benchmarks without demonstrating that they prevent real privacy or security failures.
8 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.