Research
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
Overview Research area: Multi-agent LLM systems, AI alignment, and computational moral reasoning — specifically how the structure of agent-to-agent interaction shapes the values and verdicts models pr

- arXiv
- 2510.10002
- Published
- 2025-10-11
- Authors
- Pratik S. Sachdeva, Tom van Nuenen
AI summary
Overview
- Research area: Multi-agent LLM systems, AI alignment, and computational moral reasoning — specifically how the structure of agent-to-agent interaction shapes the values and verdicts models produce.
- Technical level: Intermediate. The paper uses straightforward debate setups and a multinomial statistical model, but assumes familiarity with LLM orchestration, prompt design, and basic statistical modeling.
- Scope: Three frontier models (GPT-4.1, Claude 3.7 Sonnet, Gemini 2.0 Flash) debate blame in 1,000 contested everyday dilemmas under two orchestration protocols, revealing that protocol choice — not just model identity — drives moral judgment.
What This Paper Is About
Multi-agent debate is usually studied as a way to improve accuracy on tasks with verifiable answers. This paper asks what happens when the same debate machinery is pointed at subjective, contested decisions with no ground truth, where agents assign blame and justify moral judgments. The authors compare two ways of orchestrating that debate — models answering in parallel versus models answering in sequence — to test whether the interaction protocol itself changes what the agents conclude and which values they invoke.
Key Contributions
- A large-scale debate corpus. The authors ran 15,000 debates across protocol formats and model pairings on 1,000 dilemmas, plus an additional 15,000 debates using open-source models (DeepSeek-V3.2, Llama 3.1 8B and 70B), for more than 30,000 debates total.
- A value-level analysis of moral reasoning. Using a 48-value subset derived from Huang et al.'s Values in the Wild taxonomy (originally built from 276 second-tier values and over 3,000 empirically derived AI values), the paper traces which values each model invokes and which values a model "inherits" after changing its verdict.
- A quantitative decomposition of protocol effects. A multinomial model separates each model's inertia (tendency to repeat its prior verdict) from its conformity (tendency to follow verdicts already expressed by others), isolating order effects in round-robin debate.
- Steering experiments. System-prompt ablations test whether the observed dynamics can be redirected toward consensus-seeking, adversarial argumentation, or specific value emphasis.
Main Findings
- Models differ sharply in how often they revise verdicts. In synchronous debate, GPT-4.1 showed stronger inertia, with revision (change-of-verdict) rates of 0.6–3.1%, while Claude 3.7 Sonnet and Gemini 2.0 Flash revised at 28–41%. GPT changed its verdict in only six debates against Gemini.
- Round-robin debate flips GPT-4.1's behavior. Under sequential ordering, GPT-4.1 and Gemini 2.0 Flash emerged as highly conforming relative to Claude 3.7 Sonnet, with verdict behavior strongly shaped by order effects.
- Order effects are large. A Claude-first versus GPT debate ended in one round nearly 90% of the time, versus only 40% for the reverse order. Gemini debates ended in one round roughly 90% of the time when Gemini went second.
- Round-robin raises consensus. Three-way round-robin debate achieved consensus in virtually all dilemmas. GPT steered over 70% of dilemmas toward "NTA" when it went first and Claude third, but this effect disappeared when Claude went second. GPT's revision rate significantly increased when its position fell right after Claude.
- The statistical model quantifies the asymmetry. GPT-4.1 had the largest inertia (odds ratio 8.31; estimate 2.12, 95% CI [2.01, 2.28]), versus Claude 3.7 Sonnet at 4.39 (1.48, [1.41, 1.55]) and Gemini 2.0 Flash at 2.76 (1.02, [0.93, 1.10]). Within-round conformity was highest for GPT-4.1 (odds ratio 9.20; 2.22, [2.08, 2.34]), followed by Gemini 2.0 Flash (5.21; 1.65, [1.54, 1.79]), while Claude 3.7 Sonnet showed effectively none (1.03; 0.03, [−0.01, 0.06]). Claude was the most conforming with respect to previous rounds (1.47; 0.39, [0.35, 0.43]).
- Consensus tracks value alignment. Average value similarity (Jaccard index between the value sets each model invoked) was significantly higher when models agreed on a verdict, at roughly 0.4 to 0.5 — about three of five shared values. In debates that reached consensus, value similarity increased by 30–60%; in debates that did not, it increased only 6–17%, with mild significance observed only for GPT vs. Gemini.
- Different models emphasize different values. GPT-4.1 emphasized personal autonomy and honest communication relative to its debate partners, while Claude 3.7 Sonnet and Gemini 2.0 Flash prioritized empathetic dialogue. Specifically, GPT used Consent and personal boundaries roughly 17% more often than Gemini, while Claude favored Constructive dialogue, Conflict resolution and reconciliation, and Emotional intelligence and regulation. Claude and Gemini frequently inherited GPT's personal liberty values, while GPT most often inherited Empathy and understanding, and showed no statistically significant value inheritance from Gemini.
- Verdict distributions differ before debate. GPT heavily favored NTA (78.8% and 84.9% in its two debates), while Claude (55.6%, 55.4%) and Gemini (51.9%, 50.9%) assigned it less often. Gemini leaned on YTA (33.1%, 35.2%) far more than the others.
- Synchronous consensus rates varied by pairing. Immediate Round 1 agreement was highest for Claude vs. GPT at 66.1%, followed by GPT vs. Gemini at 53.6% and Claude vs. Gemini at 53.0%. Debates requiring additional rounds were 24.5% (Claude vs. GPT), 38.5% (Claude vs. Gemini), and 29.0% (GPT vs. Gemini); debates never converging were 9.4%, 11.5%, and 17.4% respectively.
- Steering shifts behavior but not relative dispositions. Removing the "goals" section of the system prompt had no effect on revision rates. A "balanced" prompt weighting consensus equally with correctness increased revision rates across all models (most dramatically GPT) but did not raise consensus rates. An "adversarial" prompt decreased revision rates and dropped consensus rates significantly. Prompts emphasizing Empathy and Understanding increased all models' usage of that value by 20–40% while leaving debate structure largely unchanged.
- Open-source models behave differently. DeepSeek-V3.2 resembled GPT-4.1 with low revision rates and inertia α = 2.29 and a heavy NTA emphasis, but had low conformity (γ_prev = −0.126, γ_within = −0.698). Llama 3.1 8B failed to reach consensus more often than all other models (28–31%), roughly twice as often as 70B (8–15%), despite having the highest revision rate of all models (45%) — because 8B frequently changed verdicts in debates that still failed to converge (21% of such cases, versus 5–8% for other models).
Methodology in Plain English
The authors drew on 3,272 posts and comments from Reddit's r/AmItheAsshole collected via the Reddit API for January 1 to March 31, 2025, filtering out meta, deleted, or very short posts. From these they selected the 1,000 posts with the highest commenter disagreement, on the logic that contested dilemmas stress-test value robustness. The community's five verdict categories — YTA, NTA, NAH, ESH, and INFO — became the decision space.
Using the autogen package, they ran debates in two formats. In synchronous debate, all models answer independently at the same time, then see each other's responses and may revise. In round-robin debate, models answer in sequence so that each one sees the verdicts of those who answered before it in that round. Either format runs until consensus or a maximum of four rounds.
To study values, they used Gemini 2.5 Flash as a classifier to tag up to five morally relevant values per explanation from the 48-value set, then measured similarity between two models' value sets with the Jaccard index. They validated the classifier against a human judge (100 samples, similarity 0.53), against itself across seeds and temperatures (1,602 samples, 0.64), against GPT-5 (1,602 samples, 0.55), and against a non-participating model, Kimi K2.5 (1,602 samples, 0.53) — roughly three to four overlapping values between judges.
Finally, they fit a multinomial model to debate outcomes with terms for each model's baseline verdict preference, a fixed effect per dilemma, an inertia parameter α for repeating one's own prior verdict, and conformity parameters γ for following verdicts seen in previous rounds or within the current round.
Why This Matters
Research impact. The paper reframes multi-agent debate from an accuracy-boosting technique into a variable that itself shapes model values and judgments. It introduces inertia and conformity as measurable, protocol-dependent quantities, and shows that consensus can arise from social pressure as much as from reasoning — a caution for work that treats agreement as evidence of correctness.
Real-world applications:
- Arbitration and dispute resolution, where agents render opinions or assign blame in contested cases.
- Mental health support and psychiatric assessment, where multi-agent systems are already being explored.
- Automated science systems that coordinate agents through simulated debate and critique.
- Commercial tools using multi-agent debate for market research, deal screening, and simulated focus groups.
Industry relevance. Multi-agent orchestration is increasingly a product decision, not just a research one. The finding that simply switching from parallel to sequential response ordering can turn the most rigid model into the most conforming one means the choice of protocol is a substantive design lever — one that determines whether a system's output reflects deliberation or the order in which agents happened to speak.
Future Directions
- Disentangle the drivers. The paper suggests the observed inertia and conformity likely arise from an interplay of model capacity, alignment training, and protocol, and calls for work separating these factors.
- Move beyond simple protocols. The current setup uses up to three models of similar capability; real systems may involve dozens of models with varying capabilities, hierarchical structures, independent judges, and multiple instances of the same model.
- Test alternative value frameworks. The analysis relies on a single taxonomy derived from human-Claude interactions, which may bias which value patterns appear salient; comparing against other frameworks would test generality.
- Re-examine with reasoning models. The models studied were the most advanced non-reasoning models available at the time, and are already outdated by newer reasoning-based releases, so it is unclear whether the findings generalize.
Target Audience
Researchers and practitioners working on multi-agent LLM systems, AI alignment, and AI ethics; engineers designing debate-based or judge-based orchestration for products in subjective domains; and social scientists or HCI researchers studying how AI systems negotiate values in contested, real-world decisions. The paper is accessible to readers with a working knowledge of LLMs and basic statistics.
Authors’ abstract
As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent interactions shape behavior becomes essential. Multi-agent debate has been studied as a mechanism to improve accuracy, but less is known about how debate structure -- the interaction protocol -- affects the values, dynamics, and consensus patterns that emerge when models navigate contested, real-world decisions. We address this gap by facilitating multi-agent debates among three models (GPT-4.1, Claude 3.7 Sonnet, and Gemini 2.0 Flash) to collectively assign blame in 1,000 everyday dilemmas from Reddit's ``Am I the Asshole'' community. We compare synchronous (parallel) and round-robin (sequential) interaction protocols, mirroring two fundamental ways multi-agent systems are orchestrated in practice. Across more than 30,000 total debates, our findings show striking behavioral differences, which we characterize through two dynamics: inertia and conformity. In the synchronous setting, GPT-4.1 showed stronger inertia (0.6-3.1% revision rates) than either Claude 3.7 Sonnet or Gemini 2.0 Flash (28-41% revision rates). Meanwhile, in round-robin debates, GPT-4.1 and Gemini 2.0 Flash stood out as highly conforming relative to Claude 3.7 Sonnet, with their verdict behavior strongly shaped by order effects. We further characterized the values invoked during debate, finding that GPT-4.1 emphasized personal autonomy and honest communication relative to its debate partners, while Claude 3.7 Sonnet and Gemini 2.0 Flash prioritized empathetic dialogue. Together, these results show how interaction protocol shapes moral reasoning in multi-turn debates, establishing it as a substantive sociotechnical design consideration in multi-agent systems.