Skip to content
AI.info

Research

Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations

Overview Research area: Multi-agent systems, generative agent-based modeling, and computational social science — specifically the simulation of coordinated information operations (IOs) by LLM-powered

arXiv
2510.25003
Published
2025-10-28
Authors
Gian Marco Orlando, Jinyi Ye, Valerio La Gatta, Mahdi Saeedi, Vincenzo Moscato, Emilio Ferrara, Luca Luceri

AI summary

Overview

Research area: Multi-agent systems, generative agent-based modeling, and computational social science — specifically the simulation of coordinated information operations (IOs) by LLM-powered agents on social media.

Technical level: Intermediate. The paper is readable without deep machine learning background, but it assumes some familiarity with LLM agents, network metrics (density, clustering, reciprocity), and diffusion concepts (cascades, exposure).

Scope: A systematic simulation study in which 50 LLM-powered agents (10 IO actors and 40 organic users) interact over 50 timesteps across three progressively structured operational regimes, to test whether coordination characteristic of real-world influence campaigns emerges without human guidance.

What This Paper Is About

Information operations are influence campaigns that try to manipulate public opinion on social media, and they have historically been run by human operators using manually crafted, relatively simple tactics. As LLM-based "generative agents" become more capable, there is a pressing question of whether such agents could autonomously self-organize into coordinated influence campaigns with minimal or no human oversight. The paper builds a simulated social media environment and measures whether coordination arises naturally among networked generatively agents as they are given progressively more operational awareness — from merely sharing a goal, to knowing who their teammates are, to deliberating and voting on strategy.

Key Contributions

  1. First systematic study of emergent coordination among generative agents in simulated information operations. The authors instantiate IO and organic agents in a shared simulated environment and evaluate coordination across operational regimes, rather than assuming coordination is predefined.
  2. A three-level operational regime design — Common Goal, Teammate Awareness, and Collective Decision-Making — that isolates how much coordination comes purely from shared goals versus from knowing teammates versus from explicit deliberation and voting.
  3. A dual metric framework: coordination metrics (network density, clustering coefficient, reciprocity, content similarity, comment sentiment, co-retweet similarity) and impact metrics (hashtag prevalence and adoption, time lag to adoption, exposure before adoption, engagement counts, audience diversity via the Gini coefficient, and cascade size/depth/breadth).
  4. An interactive dashboard and open code release, enabling real-time exploration of the evolving social graph, comment and re-share networks, agent-generated content, and the trajectory of hashtag adoption among organic agents.

Main Findings

  • Network cohesion increases with operational awareness (H1). The average proportion of re-shares targeting IO peers rose from 0.82 in Common Goal to 0.96 in Teammate Awareness and 0.94 in Collective Decision-Making. The comment network rose from 0.48 to 0.63 and 0.65, and the follow network from 0.27 to 0.35. Mean network density increased from 0.74 (± 0.05) to 0.89 (± 0.03) and 0.89 (± 0.02); clustering coefficient from 0.86 (± 0.01) to 0.96 (± 0.02) and 0.97 (± 0.03); reciprocity from 0.56 (± 0.07) to 0.68 (± 0.07) and 0.65 (± 0.07).

  • Narrative and affective convergence strengthen (H2). Textual similarity across IO agents' original posts increased from 0.89 (± 0.07) to 0.90 (± 0.05) and 0.91 (± 0.06), significantly higher than the organic baselines of 0.62, 0.63, and 0.61. Average sentiment of intra-group comments rose from 0.68 (± 0.03) to 0.79 (± 0.02) and 0.83 (± 0.02), also above the organic baselines of 0.62, 0.64, and 0.64.

  • Amplification becomes more synchronized (H3). Co-retweet similarity grew from 0.28 (± 0.04) to 0.31 (± 0.05) and 0.35 (± 0.03), against an organic baseline of 0.11 in all three settings. The difference between Teammate Awareness and Collective Decision-Making was also significant (p < 0.05).

  • Hashtag prevalence and adoption accelerate (H4). The proportion of original posts containing the campaign hashtag grew from 0.42 (± 0.19) to 0.47 (± 0.06) and 0.54 (± 0.01); re-shares from 0.40 (± 0.08) to 0.44 (± 0.02) and 0.47 (± 0.04); comments showed the smallest variation, 0.20 (± 0.06) to 0.20 (± 0.04) and 0.23 (± 0.01). Aligned organic agents adopted faster and more extensively than not-aligned agents in all scenarios, and the Collective Decision-Making regime produced the steepest adoption curve for not-aligned agents.

  • Adoption timing differs by political alignment (H4). Aligned organic agents showed a sharply peaked distribution near zero time lag between first IO interaction and first hashtag adoption, whereas not-aligned agents showed a broader, right-skewed distribution indicating slower and more variable uptake. Over 80% of aligned agents adopted the hashtag after exposure to only 10 IO-generated posts, while not-aligned agents required substantially more exposures.

  • Cross-group diffusion expands, but audience diversity does not (H5). Average re-shares per IO post rose from 0.75 to 1.02 and 1.19, while comments per post stayed nearly constant at 0.33, 0.34, and 0.33. Mean audience diversity scores were similar across settings: 0.624, 0.616, and 0.613, with no statistically significant pairwise differences (p > 0.05).

  • Cascades grow larger and wider (H5). Average cascade size rose from 3.84 (± 0.31) to 4.26 (± 0.20) and 4.56 (± 0.43); depth from 0.53 (± 0.02) to 0.60 (± 0.02) and 0.57 (± 0.04); breadth from 2.71 (± 0.13) to 3.08 (± 0.24) and 3.24 (± 0.31). Pairwise Mann–Whitney U tests confirmed significantly larger size and breadth in both higher-awareness settings versus Common Goal (p < 0.05), and significantly higher depth under Teammate Awareness.

  • The headline result is a low threshold for coordination. Simply revealing to agents which other agents share their goals can produce coordination levels nearly equivalent to those achieved through explicit deliberation and collective voting. In several metrics, Teammate Awareness and Collective Decision-Making produce very close values.

  • No human guidance was used. In none of the three settings were agents guided by humans in selecting actions, nor were they given explicit coordination guidelines; the regimes only modulated the information available to them.

Methodology in Plain English

The authors use Generative Agent-Based Modeling (GABM), building on an existing agent-based simulation framework (Ferraro et al., 2024) that models a Twitter/X-like platform. Each agent has three parts: a persona encoding identity and group affiliation, a memory storing interaction history, and an action policy that decides whether to post, comment, re-share, or follow by combining persona preferences with environmental feedback. The environment includes an evolving follow network and a recommender system that regulates who sees what.

Each run involves 50 agents: 10 IO actors and 40 organic agents. The 40 organic agents are split evenly into 20 aligned and 20 not-aligned with the campaign's political perspective, and their profiles are initialized from the U.S. 2020 Election dataset (Chen et al., 2022) using annotations from Ferraro et al. (2024). IO agents are prompted to promote a political candidate and maximize adoption of a campaign-specific hashtag known only inside the IO group. Each simulation runs for 50 iterations, with three repetitions per configuration.

The three operational regimes are:

  • Common Goal — IO agents know only the shared objective, not who their teammates are.
  • Teammate Awareness — agents are told the identities of their IO partners while retaining individual tactical autonomy.
  • Collective Decision-Making — every five timesteps all IO agents enter a private discussion channel with performance summaries from the previous window. Each proposes three recommendations, which an independent "IO Orchestrator" agent consolidates by identifying recurring themes, quantifying their frequency, and ranking the top five actionable strategies, which are then shared back to all IO agents. This is inspired by the Reflection Module in Park et al. (2023).

Measurements map to five hypotheses: H1 network cohesion (density, clustering, reciprocity); H2 narrative convergence (pairwise cosine similarity of posts using Sentence-BERT embeddings, plus sentiment via a RoBERTa-based classifier fine-tuned on social media text); H3 amplification (co-retweet similarity over TF-IDF vectors on a bipartite agent-to-post graph); H4 hashtag adoption (prevalence, time lag, exposures before adoption); and H5 cross-group diffusion (engagement counts, audience diversity as 1 − Gini, and cascade size/depth/breadth).

Why This Matters

Impact on research. The paper shifts the study of AI-driven influence operations from speculation to a measurable, reproducible simulation with process-level metrics rather than just outcome counts. It shows that coordination in LLM agent populations can be emergent rather than scripted, and it provides an open framework other researchers can extend. It also connects agent-based modeling to empirical IO scholarship by grounding hypotheses in documented real-world tactics such as synchronized posting, hashtag flooding, re-share rings, and coordinated reply attacks.

Real-world applications:

  • Platform governance and trust-and-safety — the coordination metrics (network density, clustering, reciprocity, co-retweet similarity) can serve as candidate early-warning indicators for detecting coordinated inauthentic behavior.
  • Election and crisis integrity monitoring — the hashtag adoption and cascade metrics point to signals that distinguish organic spread from coordinated amplification, including adoption time lag and exposures before adoption.
  • Recommender system design — the finding that negative Δt values occur when agents adopt a hashtag from their feed before any direct interaction highlights how recommendation feeds can accelerate campaign diffusion.
  • Red-teaming and policy simulation — the regime framework offers a controlled testbed for evaluating the likely effectiveness of interventions before deploying them.

Industry relevance. Social media platforms, content moderation vendors, and AI safety teams can use this work to anticipate how automated, self-organizing influence campaigns might behave at scale and to design detection and mitigation strategies against an adaptive adversary rather than a static one.

Future Directions

  • Scaling the simulation — the current setup uses only 50 agents and three repetitions per configuration, which the authors note reduces statistical power and limits reliability of significance testing for several metrics; larger populations and more runs would strengthen the evidence.
  • More realistic platform environments — extending beyond the current recommender system, memory updates, and environment feedback (detailed in Appendix B) to richer platform dynamics and cross-platform campaigns.
  • Detection and mitigation — using the identified coordination and impact signals as the basis for defenses, and testing whether such defenses hold when IO agents adapt their strategies.
  • Hybrid human-AI campaigns — real-world IOs involve both human- and automated-controlled accounts; how human operators combine with generative agents remains an open question.
  • Understanding the low-threshold effect — why simply revealing teammate identities produces coordination nearly as strong as explicit deliberation and voting is a mechanism the paper identifies but leaves for further investigation.

Target Audience

This paper is most useful to researchers in multi-agent systems, computational social science, and AI safety; platform trust-and-safety and integrity teams; and policy analysts concerned with automated influence operations. It is also relevant to practitioners building LLM agent simulations who want a worked example of measuring emergent collective behavior with network, content, and diffusion metrics.

Authors’ abstract

Generative agents are rapidly advancing in sophistication, raising urgent questions about how they might coordinate when deployed in online ecosystems. This is particularly consequential in information operations (IOs), influence campaigns that aim to manipulate public opinion on social media. While traditional IOs have been orchestrated by human operators and relied on manually crafted tactics, agentic AI promises to make campaigns more automated, adaptive, and difficult to detect. This work presents the first systematic study of emergent coordination among generative agents in simulated IO campaigns. Using generative agent-based modeling, we instantiate IO and organic agents in a simulated environment and evaluate coordination across operational regimes, from simple goal alignment to team knowledge and collective decision-making. As operational regimes become more structured, IO networks become denser and more clustered, interactions more reciprocal and positive, narratives more homogeneous, amplification more synchronized, and hashtag adoption faster and more sustained. Remarkably, simply revealing to agents which other agents share their goals can produce coordination levels nearly equivalent to those achieved through explicit deliberation and collective voting. Overall, we show that generative agents, even without human guidance, can reproduce coordination strategies characteristic of real-world IOs, underscoring the societal risks posed by increasingly automated, self-organizing IOs.

Read the original paper