Skip to content
AI.info

Research

Emergent Convergence in Multi-Agent LLM Annotation

Overview Research area: Natural Language Processing; multi-agent large language model coordination, interpretability, and automated qualitative data analysis. Technical level: Intermediate. Scope: Thi

Emergent Convergence in Multi-Agent LLM Annotation
arXiv
2512.00047
Published
2025-11-17
Authors
Angelina Parfenova, Alexander Denzler, Juergen Pfeffer

AI summary

Overview

  • Research area: Natural Language Processing; multi-agent large language model coordination, interpretability, and automated qualitative data analysis.
  • Technical level: Intermediate.
  • Scope: This paper simulates 7,500 multi-agent, multi-round LLM discussions for inductive qualitative coding and analyzes the resulting 125,000-plus utterances at the output level to show how LLM groups converge lexically, semantically, and geometrically over rounds.

What This Paper Is About

LLMs are increasingly used together in teams, but it is unclear how they coordinate when treated as black boxes. Existing work has largely studied single-turn or task-specific interactions, leaving open how models behave in sustained, group-based annotation tasks such as qualitative coding. This paper builds a large-scale simulation of multi-agent, multi-round discussions where each model proposes and revises qualitative codes, then analyzes only the resulting outputs to characterize how convergence, influence, and confidence emerge without explicit role prompting or finetuning.

Key Contributions

  1. A simulation framework for large-scale, black-box multi-agent LLM discussions applied to collaborative qualitative annotation, spanning group sizes of 2, 3, or 5 agents and discussion depths of 1–5 rounds.
  2. A set of coordination metrics capturing surface-level stability, semantic consistency, and opinion–confidence alignment, including code stability, self-consistency score, and a lexical confidence proxy.
  3. A geometric analysis of output embeddings using intrinsic dimensionality (TwoNN-Id) as a proxy for semantic compression, distinct from the internal-probe approach used in prior interpretability work.
  4. Empirical demonstration that multi-round interactions enhance lexical convergence (ROUGE), reduce embedding-space dimensionality, and produce asymmetric influence patterns between models.

Main Findings

  • Lexical convergence grows with rounds: ROUGE-1, ROUGE-2, and ROUGE-L scores between models steadily increase across successive rounds in all configurations, with the largest gains between the penultimate and final rounds. Improvements plateau after the fourth round.
  • Prompt framing matters greatly: Peak convergence occurs in the 3-model, 4-round setting for Prompt 1 (Max ROUGE-L 0.8070). Prompt 3, the academic-style "social scientist" instruction, produces the lowest ROUGE convergence of all prompt types, which the authors attribute to interpretive ambiguity and abstraction hindering lexical alignment.
  • Confidence rises overall, with model differences: Estimated from lexical certainty and hedging cues, all models increase in confidence over rounds. Mistral produces the most assertive outputs and Deepseek the most hedging-prone; the authors note Deepseek's lower scores may be partly due to its <think> reasoning content being left in during preprocessing. The lexicons contained 65 certainty expressions and 70 hedging phrases.
  • Toxicity declines: Average toxicity, measured with the Unitary Toxicity classifier, generally decreases over time. Mistral and Gemma converge to near-zero toxicity by Round 4, while Deepseek maintains relatively higher levels.
  • Stability and self-consistency vary by model: Four models maintain high stability throughout, while Deepseek shows greater variability. Deepseek and Mistral achieve the highest self-consistency, while Maverick behaves more exploratorily before converging in later rounds.
  • Asymmetric influence emerges: Early rounds show diffuse influence with Gemma and Deepseek acting as semantic anchors. Mid-discussion, Llama3.3 becomes a stronger source of influence, particularly for Gemma and Mistral, while Deepseek increasingly absorbs content from peers and acts as a semantic integrator. Later rounds (described in the paper as Rounds 6–7) show rising self- and cross-influence across all models.
  • Semantic compression in embedding space: Intrinsic dimensionality falls sharply in the 3- and 5-model setups. The 3-model group drops from an initial Id of 7.94 to 0.64 (Δ −7.30) and the 5-model group from 7.66 to 0.42 (Δ −7.24), with the steepest drops at R1. The 2-model setup stays relatively stable (13.55 to 13.11, Δ −0.44).
  • Compression is not just lexical alignment: Average pairwise cosine similarity rises only slightly overall, while intrinsic dimensionality decreases much more sharply, suggesting cosine similarity captures alignment in wording while Id reflects deeper compression of semantic space.
  • Affective and structural breakdown tracks poor convergence: Under the lowest-performing prompt, trust, joy, and valence drop sharply mid-discussion while fear, sadness, and arousal rise, with Maverick showing an affective collapse. Lexical metrics (Yule's K, hapax rate, HDD) and readability fluctuate erratically. The highest-performing prompt shows steady or improving trends in trust, dominance, positive sentiment, lexical diversity, and syntactic depth.
  • Convergence can help or harm quality: A semantic-flattening case saw diverse initial codes collapse into "Exasperated Urban Compassion Fatigue & Policy Critique" with ROUGE-L rising by +0.61 and intrinsic dimensionality falling from 6.10 to 1.07, risking erasure of sub-themes. A positive case produced consensus on "Challenging Sexist Stereotypes in Media," raising ROUGE-L by +0.45 and cosine similarity by +0.24 while preserving meaning.
  • Compensatory grounding: Under the worst-performing prompt, several models (e.g., Mistral, Maverick) show stable or increasing visual sensorimotor activation despite declining socialness, concreteness, and emotional coherence, which the authors read as a possible fallback to concrete language when abstract coordination fails.

Methodology in Plain English

The researchers built a synthetic annotation team and watched what happened. They drew 500 English comments from the Jigsaw Unintended Bias in Toxicity Classification dataset, choosing comments with high annotator disagreement (to capture subjectivity) and a minimum length of 100 words (to ensure interpretive richness).

Each simulated discussion ran in three phases: every agent generated an initial code, then agents took turns refining codes over one to five rounds in a fixed order, and finally each agent produced a synthesis code. After every turn, an agent summarized its message in a single sentence; these summaries accumulated into a shared conversational memory that served as context for later turns. This avoided external memory modules and finetuning, relying on prompt-based inference alone, and required no explicit roles for the agents.

The team varied group size (2, 3, or 5 agents) and discussion depth (1–5 rounds), generating 500 discussions per configuration for 7,500 total simulations. Five prompt templates, ranging from formal thematic-analysis instructions to informal summary requests, were rotated across discussions for balanced coverage.

Analysis happened entirely at the output level. The authors measured lexical convergence with ROUGE-1, ROUGE-2, and ROUGE-L; projected sentence embeddings with UMAP; scored toxicity with the Unitary Toxicity classifier; and extracted 190 discussion-level linguistic features with the ELFEN toolkit covering syntax, lexical diversity, readability, emotion, and psycholinguistic norms. They added three process metrics: code stability (proportion of string-identical outputs between consecutive rounds), a self-consistency score (cosine similarity of TF-IDF representations between rounds), and a confidence score (the length-normalized difference between certainty and hedging cue counts). For geometry, they estimated intrinsic dimensionality with the Two-Nearest Neighbor method (TwoNN-Id) on external sentence-transformer embeddings, treating it as a proxy for the semantic complexity of outputs rather than a probe of internal states.

Why This Matters

Research impact. The paper offers a scalable, output-only complement to probe-based interpretability. Instead of inspecting hidden states, it shows that observable interaction patterns—convergence, influence asymmetry, dimensionality shrinkage—can reveal coordination dynamics. It also connects multi-agent LLM behavior to decades of social psychology work on group decision-making, notably the opinion–confidence mapping of Moussaïd et al. (2013), and extends qualitative-coding research from single-model annotation to team-based settings.

Real-world applications:

  • Collaborative qualitative coding of open-ended text such as survey responses, interview transcripts, or social media comments.
  • Content moderation and bias assessment, given the toxicity trends and the use of the Jigsaw unintended-bias dataset.
  • Decision-support systems that need documented consensus among multiple model perspectives.
  • Human–LLM hybrid annotation teams, where the paper's metrics could flag when models are over-converging and erasing nuance.

Industry relevance. Organizations deploying LLM ensembles for labeling or analysis can use these metrics to monitor whether a group is producing a genuinely shared understanding or merely collapsing into uniform phrasing. The finding that convergence can destroy useful sub-themes is directly actionable: the authors recommend monitoring for semantic drift and incorporating occasional human oversight. The asymmetric influence patterns also suggest that model composition choices matter as much as individual model quality.

Future Directions

  1. Extending the framework to mixed human–LLM teams, where human judgment could counteract semantic flattening.
  2. Testing whether consensus remains stable under noisy or adversarial conditions.
  3. Analyzing how turn-taking order, memory design, and agent identity affect group outcomes.
  4. Replacing the lexical confidence proxy with learned measures or alternative normalizations (per word or per sentence), and testing whether findings generalize beyond the single dataset and five-prompt setup used here.

Target Audience

Researchers and practitioners in NLP interpretability, multi-agent systems, and computational social science will get the most from this paper, along with qualitative researchers exploring LLM-assisted coding. It also suits applied teams building annotation or consensus pipelines who need concrete metrics for monitoring group convergence and detecting when agreement comes at the cost of nuance. Readers should be comfortable with metrics such as ROUGE, cosine similarity, TF-IDF, and embedding dimensionality.

Authors’ abstract

Large language models (LLMs) are increasingly deployed in collaborative settings, yet little is known about how they coordinate when treated as black-box agents. We simulate 7500 multi-agent, multi-round discussions in an inductive coding task, generating over 125000 utterances that capture both final annotations and their interactional histories. We introduce process-level metrics: code stability, semantic self-consistency, and lexical confidence alongside sentiment and convergence measures, to track coordination dynamics. To probe deeper alignment signals, we analyze the evolving geometry of output embeddings, showing that intrinsic dimensionality declines over rounds, suggesting semantic compression. The results reveal that LLM groups converge lexically and semantically, develop asymmetric influence patterns, and exhibit negotiation-like behaviors despite the absence of explicit role prompting. This work demonstrates how black-box interaction analysis can surface emergent coordination strategies, offering a scalable complement to internal probe-based interpretability methods.

Read the original paper