Social Atoms / Research

Awesome Social Simulation

Simulating people and societies with language models.

14 papers5 topics2022–2026

A map of social simulation

Explore social simulation by topic Individuals, interactions, and societies are connected scales of social simulation. Choose one of five reading entry points: silicon participants, social science experiments, multi-agent interaction, control and calibration, or evaluation. Human evidence grounds work across every scale. Individuals Silicon participants Interactions Social science experiments Societies Multi-agent interaction Human evidence Control and calibration Evaluation
Explore social simulation by topic Individuals, interactions, and societies are connected scales of social simulation. Choose one of five reading entry points: silicon participants, social science experiments, multi-agent interaction, control and calibration, or evaluation. Human evidence grounds work across every scale. Individuals Silicon participants Interactions Social science experiments Societies Multi-agent interaction Human evidence across every scale Control and calibration Evaluation

Select a tag in the map to filter papers.

About the map & further reading

Individuals, interactions, and societies influence each other; evidence, calibration, and evaluation matter at every scale. Tags are reading entry points and can span several scales. The map offers an orientation to the field, informed by the readings below.

  1. From Individual to SocietyMou et al. · 2024 · A survey across individual, scenario, and society-level simulation.
  2. Large language models empowered agent-based modeling and simulationGao et al. · 2024 · Agents, environments, interactions, and their design challenges.
  3. Validation is the central challenge for generative social simulationLarooij & Törnberg · 2026 · A critical review of validation and the claims simulations can support.

All papers

14 papers

Browse the full collection, or follow a topic.

HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning

EMNLP 2026 Main · AcceptedBenchmark

Authors & publication details

Chance Jiajie Li*, Zhenze Mo*, Yuhan Tang*, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson

* Equal contribution.

Accepted to EMNLP 2026 Main. Earlier presentations: NeurIPS 2025 workshops — PersonaLLM (Oral) and LAW (Spotlight). Publication history ↗

Benchmarks whether models can recover a person’s beliefs and predict belief updates under intervention, using questionnaires and think-aloud interviews across three policy domains.

Notes across topics
  • Evaluation

    Benchmarks whether models can recover a person’s beliefs and predict belief updates under intervention, using questionnaires and think-aloud interviews across three policy domains.

  • Silicon participants

    Tests individual-level reasoning simulation using human participants’ self-reported beliefs and reasoning traces, rather than population-level averages.

Simulating Society Requires Simulating Thought

NeurIPS 2025 · Position TrackPosition paper

Authors & publication details

Chance Jiajie Li*, Jiayi Wu*, Zhenze Mo, Ao Qu, Yuhan Tang, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Jinhua Zhao, Paul Liang, Luis Alonso, Kent Larson

* Equal contribution.

Published in the NeurIPS 2025 Position Paper Track.

Proposes GenMinds: cognitively grounded agents with structured, revisable beliefs and traceable reasoning for social simulation.

Notes across topics
  • Control and calibration

    Proposes GenMinds: cognitively grounded agents with structured, revisable beliefs and traceable reasoning for social simulation.

  • Evaluation

    Proposes RECAP to evaluate reasoning fidelity through causal traceability, demographic grounding, and intervention consistency.

Compares language-model opinion distributions with public-opinion surveys across 60 demographic groups.

Notes across topics
  • Evaluation

    Compares language-model opinion distributions with public-opinion surveys across 60 demographic groups.

  • Silicon participants

    Measures how well model responses represent opinion distributions across demographic groups.

Generates populated online communities with diverse member personas, posts, replies, and social interventions.

Notes across topics
  • Evaluation

    Evaluates whether generated online communities and interactions resemble real community behavior and support design exploration.

  • Multi-agent interaction

    Generates populated online communities with diverse member personas, posts, replies, and social interventions.

End of collection · 14 papersBack to top ↑

A focused reading list.

Maintained by Social Atoms. Each paper is selected for its contribution to simulating human behavior, social interaction, or collective dynamics.

What belongs in the list?

Peer-reviewed papers published or accepted at established conferences or journals. Original research, benchmarks, and selected position papers are included, with publication track and acceptance status shown explicitly.

The collection excludes stand-alone preprints, workshop-only entries, and surveys. Earlier workshop presentations may appear in a paper’s publication history. General-purpose multi-agent task solving is outside its scope.

Years refer to the publication or acceptance venue, not the first preprint. Topic labels describe a paper’s focus; they are not ratings of evidence or reliability.

How can I contribute?

Suggest a paper with its official publication link, venue, and a short explanation of its social-simulation contribution. Use the existing topics and prioritize the paper’s main contribution.

Contribute on GitHub ↗