Skip to content
AI.info

Research

On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models

Summary: On Emergent Social World Models Overview Research area: Mechanistic interpretability and cognitive-science-inspired evaluation of large language models, at the intersection of Theory of Mind

On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models
arXiv
2602.10298
Published
2026-02-10
Authors
Polina Tsvilodub, Jan-Felix Klumpp, Amir Mohammadpour, Jennifer Hu, Michael Franke

AI summary

Summary: On Emergent Social World Models

Overview

Research area: Mechanistic interpretability and cognitive-science-inspired evaluation of large language models, at the intersection of Theory of Mind (ToM), pragmatics, and functional localization.

Technical level: Advanced. The paper combines behavioral benchmarking of 48 language models with causal ablation experiments on localized subnetworks and Bayesian hierarchical regression modeling.

One-sentence scope: The paper tests whether language models reuse shared internal computational mechanisms for Theory of Mind and pragmatic reasoning — the "functional integration hypothesis" — using behavioral evaluations and causal-mechanistic ablation experiments inspired by cognitive neuroscience.

What This Paper Is About

Debate persists over whether large language models genuinely develop "world models" or merely mimic surface statistics. Social reasoning is a hard test case, because the mental states that drive it (beliefs, desires, intentions) are latent and are not directly observable in text. The authors ask whether LMs' capacities for general Theory of Mind and for language-specific pragmatic reasoning draw on the same internal machinery, which would count as evidence for an emergent "social world model" whose representations are repurposed across tasks.

Key Contributions

  1. A testable framing of two competing hypotheses. The paper formalizes the functional specialization hypothesis (different mechanisms for pragmatics and ToM) against the functional integration hypothesis (shared mechanisms), and derives concrete statistical predictions from each for both behavioral and causal-mechanistic tests.

  2. Novel ToM localizer data. The authors contribute four synthetically generated localization suites built from experimental materials previously used in fMRI studies, totaling 1400 stimuli, spanning the seven ATOMS ToM subcategories. This is a substantially larger localizer dataset than used in prior like-minded work, which had not found strong evidence for functional ToM localization.

  3. Methodological refinements to functional localization for LMs. The paper ports the conjunctive "minimum statistic" approach from cognitive neuroscience into LM analysis, compares it against a simpler approach, and reports causal ablation results across 20 models and eight localizers.

  4. Fine-grained behavioral and mechanistic analysis of ToM subcategories. Using the ATOMS framework, the authors annotate evaluation datasets and run model comparisons over all combinations of the seven ToM aspects, linking behavioral findings to ablation results.

Main Findings

  • Behavioral correlation (P1) confirmed. Across 48 models, average accuracy on ToM and pragmatics tasks showed a moderate, significant positive correlation (r = 0.68, p = 1.24 × 10⁻⁷).

  • No domain effect in the regression (P2) confirmed. In a Bayesian beta regression, there was no credible difference between ToM and pragmatics domains (β = -0.03 [-0.74, 0.67]), meaning domain information does not help predict a model's accuracy once enough is known about the model. Fine-tuned models outperformed base models (β = 0.11 [0.05, 0.16]); large models outperformed medium (β = 0.20 [0.11, 0.29]) and medium outperformed small (β = 0.26 [0.20, 0.32]). No difference was found between dataset types (β = -0.06 [-0.88, 0.73]).

  • ToM accuracy predicts pragmatic accuracy (P3) confirmed. Comparing two Bayesian beta regression models via leave-one-out cross-validation, the model using ToM accuracy as a predictor outperformed the model using general linguistic benchmark accuracy (ELPD = -16.1, p-value = 0).

  • ATOMS subcategories carry predictive signal. Comparing 128 regression models on ToM accuracy, the baseline model was numerically worse than every model containing ATOMS predictors. The best model included intentions, desires, emotions, percepts and non-literal communication (ELPD = -58.80, p-value = 0). The percepts predictor added the most predictive power, distinguishing false-belief location tasks and agent-property tasks from other datasets. The weakest models were those lacking the percepts predictor.

  • Localization is reproducible. Across 10-fold cross-validation, the highest generalization rates (≥ 9 significant folds) were found for the all, MoralIntent and GameBeliefs localizers; high rates (≥ 8 folds) for the simple CommunicativeIntent and LatentBeliefs localizers; and lower rates (5–7 folds) for conjunctive suites. Results for the LB + CI suite were excluded due to a coding error.

  • Causal ablation evidence from the global analysis. Marginalizing over all models and localizers, ToM-network ablation decreased ToM performance (β = 0.25 [0.14, 0.35]) and did so more than the control ablation (β = 0.06 [0.02, 0.11]). ToM-network ablation also decreased pragmatic performance (β = 0.30 [0.20, 0.39]) more than the control (β = 0.07 [0.02, 0.12]). Ablation did not credibly decrease baseline linguistic benchmark performance (β = 0.13 [-0.10, 0.35]), consistent with the prediction that it would not, but the prediction that the causal effect would be larger for pragmatics than for baseline was not supported (β = 0.17 [-0.07, 0.41]).

  • Localizer-specific variation. Predictions 1.1, 2.1 and 3.1 held for all localizers except CommunicativeIntent (conjunctive) for 1.1 and GameBeliefs for 3.1. The differential-effect predictions 1.2 and 2.2 held only for LatentBeliefs (simple and marginally conjunctive) and GameBeliefs, suggesting these three localizers drive the global results and that the ToM aspects beliefs, percepts, desires and emotions may functionally support both ToM and pragmatic performance.

  • Entity tracking shares mechanisms. Ablating critical subnetworks localized by LatentBeliefs and GameBeliefs affected entity tracking performance similarly to ToM and pragmatic performance across models, but did not as consistently affect other datasets, suggesting ToM and pragmatic capabilities may also rely on shared entity tracking mechanisms.

  • Simple and conjunctive localization are roughly interchangeable. No credible differences were found between the two methods for LatentBeliefs (β = 0.05 [-0.03, 0.15]) or CommunicativeIntent (β = 0.06 [-0.03, 0.15]) relative to baseline, nor relative to control ablation (LB: β = 0.10 [-0.03, 0.23]; CI: β = 0.07 [-0.06, 0.20]). The proportion of units localized by both methods was LB = 0.50 and CI = 0.66.

  • Model size does not credibly change the causal effect. Averaged across localizers, no credible differences were found between large and medium models (ToM: β = 0.15 [-0.12, 0.42]; pragmatics: β = 0.13 [-0.12, 0.38]) or between medium and small models (ToM: β = -0.04 [-0.26, 0.18]; pragmatics: β = -0.11 [-0.41, 0.20]).

  • An unexplained result. The paper notes it remains puzzling why ablating the subnetwork identified by the CommunicativeIntent localizer did not have the predicted causal effect, despite that suite being considered particularly relevant for comparison with pragmatic reasoning.

Methodology in Plain English

Behavioral evaluation. The authors assembled 16 pragmatic datasets covering ten pragmatic phenomena and 22 ToM datasets (several already annotated with ATOMS subcategories by prior work, others annotated by the authors). Every item is multiple choice with between 2 and 6 answer options. They tested 48 open-weights or open-source models from seven families (Llama, Qwen, Falcon, Mistral, OLMo, Pythia, Gemma), ranging from 0.5B to 72B parameters, testing base and fine-tuned versions where available. Rather than generating text, they scored each answer option by its conditional log-probability given the prompt, averaged across tokens to correct for differing answer lengths, and took the highest-scoring option as the model's answer. General language ability was measured with BLiMP and SNLI. They then fitted Bayesian beta regression models in R using the brms package, with model family, size, type, dataset type and domain as predictors, and compared models using leave-one-out cross-validation.

Functional localization. Borrowing from cognitive neuroscience, the authors built "localizer suites" that contrast target stimuli requiring the capacity of interest against closely matched control stimuli that do not. Four suites were built on materials from prior fMRI studies: LatentBeliefs, CommunicativeIntent, GameBeliefs and MoralIntent. To reduce contamination and surface-level confounds, they used GPT-5 (which was not itself evaluated) with a 5-shot prompt to generate 100 novel synthetic stimuli per condition per suite, then used principal component analysis on SentenceTransformer all-MiniLM-L6-v2 embeddings to check that synthetic and original stimuli did not form clearly separate clusters — which held for all conditions except HumanDescr.

For each model, they recorded unit activations at the last token before answering, then identified units whose activations differed between target and control stimuli using Welch's t-test (paired for the simple suites GameBeliefs and MoralIntent, unpaired otherwise). Two localization strategies were used: a simple approach comparing the union of all target against all control stimuli, and a conjunctive approach taking the minimum t-statistic across all target–control pair comparisons, following the minimum-statistic approach from neuroscience. Units were selected at significance level α = 0.05, capped at the top 1% by absolute statistic; a control subnetwork of equally many least-active non-significant units was also constructed. This yielded eight localizers: five simple (one per suite plus a union "all" localizer) and three conjunctive (LatentBeliefs, CommunicativeIntent, and their union LB + CI).

Causal ablation. Twenty models from four families (Qwen-2.5, Llama, Falcon-3, Gemma-2) with both base and fine-tuned versions were selected. For each model, 19 to 24 datasets from the 36 behavioral datasets on which the model performed above chance were used for testing. Critically, the localized units were set to zero and performance was re-measured on ToM datasets, pragmatic datasets, and baseline linguistic benchmarks, plus entity tracking and analogical reasoning tasks in exploratory analyses.

Why This Matters

Impact on research. The paper offers a template for testing claims about LM "world models" with falsifiable, pre-specified predictions rather than loose analogies. It also supplies new localizer data and a demonstration that neuroscience methods such as conjunctive minimum-statistic analysis can be carried over to transformer models. The finding that the ATOMS percepts predictor carried the most signal, and that the features driving the best behavioral models overlapped with the subnetworks whose ablation had the strongest causal effects, points toward a more granular, subcategory-level account of social reasoning in LMs rather than a single monolithic ToM faculty.

Potential real-world applications (as implications of this line of work, not claims verified in the paper):

  • Safety and alignment auditing: locating the internal units that support social reasoning could inform targeted inspection or intervention in models deployed in socially sensitive settings.
  • Model editing and control: knowing which subnetworks causally support which social abilities could guide surgical modification rather than retraining.
  • Evaluation design: the finding that ToM and pragmatic accuracy covary across models can inform more efficient benchmark suites, since one domain may be partly predictive of the other.
  • Human–machine communication and human cognition research: the shared-mechanism framing offers a concrete way to compare human neural accounts of ToM and pragmatics with machine counterparts.

Industry relevance. The paper suggests that social reasoning ability is not a set of isolated skills bolted onto a language model but is at least partly entangled with general-purpose machinery, and that model scale does not credibly change the size of these causal effects. That has consequences for practitioners who assume that larger models will develop separable, independently steerable social competencies.

Future Directions

  • Resolve the CommunicativeIntent puzzle. The predicted causal effect failed to appear for the localizer suite judged most relevant to pragmatics, and the reasons for this are left open.

  • Disentangle social reasoning from general intelligence. Entity tracking appeared to share mechanisms with ToM and pragmatic performance, but the exact nature of the general mechanisms involved remains elusive for both humans and machines.

  • Improve grounding of the concepts. The authors note that the results hinge on whether the datasets truly operationalize general ToM, specific ATOMS aspects, and pragmatic reasoning, and that ATOMS annotations were partly author-provided, with a low number of annotations and potential noise. Crowdsourced annotation is suggested.

  • Test robustness to design choices. Ablation was implemented by setting units to zero, subnetwork size was capped at the top 1% of units, and single prompt formats were used; the paper calls for exploring alternative ablation methods and different ablation rates, extending beyond BLiMP and SNLI for linguistic baselines, and moving beyond English-only, resource-rich-language datasets.

Target Audience

Researchers in NLP interpretability and mechanistic analysis, cognitive scientists and psycholinguists studying Theory of Mind and pragmatics, and alignment or evaluation practitioners who care about how social reasoning capabilities are organized inside language models. The paper assumes familiarity with transformer internals, activation-based analysis, and Bayesian regression, so readers without that background will find the methodology sections demanding, though the motivation and conclusions are accessible.

Authors’ abstract

This paper investigates whether LMs recruit shared computational mechanisms for general Theory of Mind (ToM) and language-specific pragmatic reasoning in order to contribute to the general question of whether LMs may be said to have emergent "social world models", i.e., representations of mental states that are repurposed across tasks (the functional integration hypothesis). Using behavioral evaluations and causal-mechanistic experiments via functional localization methods inspired by cognitive neuroscience, we analyze LMs' performance across seven subcategories of ToM abilities (Beaudoin et al., 2020) on a substantially larger localizer dataset than used in prior like-minded work. Results from stringent hypothesis-driven statistical testing offer suggestive evidence for the functional integration hypothesis, indicating that LMs may develop interconnected "social world models" rather than isolated competencies. This work contributes novel ToM localizer data, methodological refinements to functional localization techniques, and empirical insights into the emergence of social cognition in artificial systems.

Read the original paper