Research
Multilingual Agent-Based World Modeling for Social Science
Overview Research area: Computational social science / LLM multi-agent simulation (Natural Language Processing, cs.CL). Technical level: Advanced. The paper assumes familiarity with LLM-based agents,
- arXiv
- 2512.07195
- Published
- 2025-12-08
- Authors
- Xuan Zhang, Wenxuan Zhang, Anxu Wang, See-Kiong Ng, Yang Deng
AI summary
Overview
Research area: Computational social science / LLM multi-agent simulation (Natural Language Processing, cs.CL).
Technical level: Advanced. The paper assumes familiarity with LLM-based agents, retrieval/recommendation systems, embedding models, and quantitative calibration metrics such as RMSE.
Scope: This paper introduces MAWM, a framework for simulating multi-turn multilingual interaction among generative agents with sociolinguistic personas, together with the MAPS benchmark built from global survey data.
What This Paper Is About
Existing LLM multi-agent simulations of society are almost entirely monolingual and culturally homogeneous, so they cannot represent the cross-lingual discourse through which real opinions form and shift across linguistic boundaries. The authors build MAWM, the first framework they describe as Multilingual Agent-based World Modeling, in which user agents and news-organization agents speaking different native languages read, write, and react to one another on a simulated social platform. The goal is to test whether native-language simulation produces attitude dynamics that better match real survey data than English-only simulation, and to use the resulting society for social science case studies.
Key Contributions
-
MAWM framework. The first multilingual agent-based world modeling framework that situates multi-agent society simulation in a global context, allowing agents to communicate across languages and cultures. It supports two analysis modes: global public opinion modeling and media influence / information diffusion via autonomous news agents that generate content conditioned on institutional profiles and evolving discourse.
-
MAPS benchmark. The Multilingual Agent Perspective Survey dataset pairs survey questions from GlobalOpinionQA (derived from the World Values Survey and Pew Global Attitudes Survey) with demographic personas derived from the World Values Survey, spanning 50 countries and 28 languages. Personas use eight independent attributes (age, education, gender, marital status, occupation, political preference, religion, social class) plus country and native language.
-
Evaluation of simulation reliability. The paper evaluates real-world calibration (RMSE against real survey distributions), global sensitivity (response to injected positive/negative news), and local consistency (LLM-judge ratings of agent action quality), across five backbone LLMs.
-
Social science case studies. Case studies on cultural assimilation (trade and domestic wages) and normative diffusion (gender equality norms) show that the simulated agent society reproduces established sociocultural phenomena — filter bubbles, asymmetric media influence, and differential convergence rates — that monolingual simulations cannot capture.
Main Findings
-
Native-language simulation calibrates better in most countries. Comparing English versus native-language simulation across all 21 non-English countries in MAPS, native-language simulation yields lower RMSE in 13 of the 21 countries, with the largest gains (up to 0.18 RMSE) concentrated in lower-resource, largely non-Latin-script languages such as Ukrainian, Thai, Arabic, and Bengali.
-
Where English wins, the margin is small. English outperforms native-language simulation only with |Δ| ≤ 0.11, almost entirely in high-resource Spanish-speaking countries, which the authors characterize as a weak test because Spanish behaves much like English.
-
Model differences in initial calibration. In the no-news, no-cross-country-communication setting, Llama-4-Maverick and Llama-4-Scout perform uniformly better than the other tested models; Llama-4-Maverick is selected as the best performer for subsequent experiments and case studies.
-
Agents respond to injected news. Global sensitivity results across all seven studied countries show that user agents shift attitudes toward the designated editorial stance under both positive and negative news exposure, measured with a baseline-normalized shift Δ̂_o = Δ_o / (1 − p̃_o).
-
News creation is the weakest agent behavior. LLM-judge scoring (GPT-5, 1 to 5 scale, 15 user agents per case plus all news agents, over 5 rounds) shows agents generally achieve high response quality except for news creation, where news agents tend to repeat previous posts and reduce content diversity over rounds.
-
Different countries converge through different mechanisms. For Case 1 (India, Japan, United States), simulated U.S. users remain stable at scores of 1.0 to 1.2, Japan converges toward India with scores stabilizing around 0.5–0.6 while its foreign exposure approaches zero by round 10. For Case 2 (South Korea, Brazil, Peru), South Korea falls from about 0.8 to about 0.3, ending more supportive of free trade, while Brazil and Peru converge within five communication rounds at about 0.4.
-
Early recommendation dominance can steer convergence. Brazilian and Peruvian content accounts for 73.3% of South Korea's foreign recommendations in round 1, declining to 3.5% by round 8, and this early surge coincides with South Korea's opinion shift.
-
Filter bubbles appear in multiple countries. Japan shows the highest inbreeding-homophily index, and Peru's foreign content exposure declines over rounds, mirroring Japan. The authors distinguish Japan's recommender-driven echo chamber from South Korea's convergence through genuine exposure to similar-stance content.
-
News organizations diffuse norms more strongly than users. In Case 3, Dutch user posts make up less than 2% of Zimbabwean agents' recommended content, whereas news organizations have larger impact because the recommendation system includes at least one news post per round; Zimbabwean agents shift in a pro-equality direction in response.
-
Normative diffusion is asymmetric. Dutch users' attitudes remain consistent throughout the simulation despite sustained exposure to 100 Zimbabwean users with opposing views, confirming that diffusion runs from source to target rather than mutually.
-
External evidence supports the direction of simulated change. In World Values Survey waves for Zimbabwe, agreement that "men make better political leaders than women" moved from a 52% majority (2001) to 58% (2012) to a 45% minority (2020), and disagreement that "a university education is more important for a boy than a girl" rose from 81% to 86%. Internet penetration is positively associated with gender equality across more than 50 countries (β = 0.0014, p < 0.01) after controlling for GDP per capita, urbanization, enrollment, and region. The authors claim only the direction and driver of change, not its exact rate.
Methodology in Plain English
The researchers build a small simulated social platform. Each run starts with 100 user agents whose personas are drawn from World Values Survey demographics, distributed across countries by population proportions, with each agent assigned a native language based on country. News organization agents are given an editorial stance and a language.
A multilingual recommendation system, using a Sentence Transformer (concretely jina-embeddings-v3) plus Google Translate, projects agents and posts of different languages into one embedding space. Posts are ranked by semantic relevance combined with recency decay, and each agent receives the top-k_r posts, excluding its own, translated into its own language.
The simulation runs a warm-up round 0, where agents write self-introductions, produce initial posts, and vote, then 20 further rounds. In later rounds, user agents read recommended posts and extract weighted takeaways into long-term memory, write new posts, and vote to update their attitude distribution over the survey answer options. Short-term memory comes from three chain-of-persona self-questioning steps before each action. News agents instead write directly from recommended content and post history, preserving asymmetry with users.
Evaluation has three parts: RMSE between simulated and real survey distributions (ranging from 0 for a perfect match to sqrt(2/|C|) for maximum divergence), the media-induced attitude shift normalized by remaining headroom, and an LLM-judge rating of action quality. Three cases are used: Q201 from the Pew Global Attitudes Survey on whether trade increases, decreases, or does not affect wages of one's nationality's workers (Cases 1 and 2), and Q278 from the World Values Survey on whether a girl should honor her family's wishes even if she does not want to marry (Case 3). Case 3 pairs Zimbabwe (ranked 153rd on the Gender Inequality Index) with the Netherlands (8th). Five LLM backbones are tested: GPT-5-mini, GPT-4o-mini, Gemini-2.5-Flash, Llama-4-Maverick, and Llama-4-Scout. Each case is run for three trials.
Why This Matters
Impact on research. The paper argues that language is not a cosmetic variable in agent simulation: prompting the same model in different languages yields measurably different cultural cognitive styles, so English-only simulation suppresses culture-specific perspectives. It reframes multilingual simulation as a controlled, scalable complement to global surveys that capture only static snapshots, and it provides an open benchmark (MAPS) plus released code and data for cross-cultural computational social science.
Real-world applications.
- Estimating cross-national public opinion on open-domain topics without fielding new multinational surveys.
- Studying how media organizations pierce or reinforce filter bubbles across language boundaries.
- Modeling information diffusion and normative change, such as gender equality norms moving from a source country to a target country.
- Testing how opinion converges or fails to converge when communities that speak different languages are placed on a shared platform.
Industry relevance. The framework mirrors deployed production systems that translate cross-language posts and rank them by predicted user interactions, so its findings about early recommendation dominance, homophily-driven echo chambers, and the outsized reach of news accounts speak directly to platform design, recommendation auditing, and content-moderation questions in multilingual markets. The released benchmark also gives model developers a way to measure whether their systems behave differently, and less accurately, when operating in lower-resource languages.
Future Directions
-
Scale beyond 100 agents. The authors explicitly note that simulations use 100 agents because of computing resource limitations, and leave scaling to larger populations and to frameworks with a user pool of millions of individuals to future work.
-
Improve persona fidelity. Demographic attributes alone have limited power in aligning LLM role-playing with real human opinions, so incorporating belief-level information into personas is flagged as future work. Persona-to-opinion mapping currently operates at the country level, which may obscure within-country variation.
-
Fix weak news generation. News agents repeat previous posts and lose content diversity over rounds, which is the lowest-scoring agent behavior in the local consistency evaluation and a clear target for improvement.
-
Address simulation and network limitations. The social network is a deliberately simplified abstraction; real communication involves imperfect translation and selective or slanted reporting. The authors also state that the current calibration and evaluation have limitations discussed in Appendix A, and that prompt templates and hyperparameters may need adaptation for specific research scenarios.
Target Audience
Researchers in computational social science and LLM-based multi-agent simulation who want a multilingual, culturally grounded environment for studying opinion dynamics; NLP researchers working on cross-cultural or multilingual model behavior; and applied practitioners in platform recommendation, media analysis, and policy research who need a controllable, interpretable alternative to large multinational surveys. Readers should be comfortable with quantitative evaluation, benchmark design, and agent architectures.
Authors’ abstract
Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly monolingual without cross-lingual interaction, an essential property of real societies. We introduce MAWM, the first Multilingual Agent-based World Modeling framework that supports multi-turn multilingual interactions among generative agents with diverse sociolinguistic profiles. MAWM enables two modes of analysis: (i) global public opinion modeling, which tracks how attitudes toward open-domain survey questions evolve across languages and cultures, and (ii) media influence and information diffusion, via autonomous news agents that dynamically generate content and shape user behavior. To ground the simulation in realistic population distributions, we construct the MAPS benchmark, which combines survey questions and demographic personas drawn from global population distributions. Experiments on simulation alignment, along with social science case studies on cultural assimilation and normative diffusion, show that native-language simulation better reflects real survey data than English-only simulation, and that the agent society in MAWM reproduces established sociocultural phenomena and highlights the value of multilingual simulation as an alternative interpretable tool for computational social science.