Skip to content
AI.info

Research

Understanding Persuasion in Long-Running Agents

Understanding Persuasion in Long-Running Agents Overview Research area: AI agent safety and reliability — specifically how persuasive natural-language input influences the downstream, multi-step behav

Understanding Persuasion in Long-Running Agents
arXiv
2602.00851
Published
2026-01-31
Authors
Hyejun Jeong, Amir Houmansadr, Shlomo Zilberstein, Eugene Bagdasarian

AI summary

Understanding Persuasion in Long-Running Agents

Overview

Research area: AI agent safety and reliability — specifically how persuasive natural-language input influences the downstream, multi-step behavior of tool-using LLM agents.

Technical level: Intermediate. The paper combines a conceptual framing from social psychology (persuasion, cognitive dissonance) with empirical agent evaluation, using behavior-trace metrics, composite scores, PCA-based constructs, and rank-normalized statistics.

Scope in one sentence: The paper introduces and empirically tests "persuasion propagation" — the idea that task-irrelevant persuasion, or an explicitly specified belief state, can shift how a long-running agent searches, codes, and revises even when its final outputs look normal.

What This Paper Is About

Modern AI agents do more than chat: they browse the web, write and debug code, and run many tool calls across extended sessions. The authors ask what happens when such an agent is exposed to a persuasive claim that is completely irrelevant to the task it is later given — for example, a stance about government regulation of online privacy, followed by a research task about quitting smoking.

Existing persuasion research in LLMs mostly measures whether a model changes its stated stance, typically in single-turn settings. This paper shifts the question to behavior: does an adopted belief persist and quietly reshape how the agent plans, searches, selects sources, and revises code? The authors name this phenomenon persuasion propagation and build a controlled framework to isolate it from ordinary instruction-following or prompt injection.

Key Contributions

  1. Problem formulation and framing. The paper defines persuasion propagation, a phenomenon in which task-irrelevant persuasion affects an agent's downstream execution behavior — distinct from prompt injection or direct instruction following, because the persuasive content is never part of the task context and may never appear in the final answer.

  2. Controlled isolation of belief effects. The authors propose an evaluation framework that separates belief state from persuasion timing, comparing on-the-fly persuasion (real-time vulnerability during tasks) against prefilled belief conditioning (belief, disbelief, or neutrality specified directly at task time).

  3. Trace-level evaluation. They show that persuasion propagation appears as behavioral drift that is not visible from final outputs alone, motivating trace-based metrics — coding scores (TRS, EVS) and web-research constructs (activity, breadth, depth) — for auditing LLM agents.

  4. Empirical evidence of behavioral drift. Across personas, persuasion tactics, and three task families, the authors report evidence that persuasion propagates across irrelevant tasks, with a task-irrelevance check (SBERT similarity mean = 0.023, std = 0.067 versus a positive-control mean = 0.110, std = 0.110) confirming minimal semantic overlap between injected claims and task prompts.

Main Findings

  • Persuasion tactics change stated belief, but not equally across backbones. On the controversial-claim task with 8 distractor interventions, gpt-4.1-nano and mistral-nemo-12b show increased persistence under every tactic relative to the no-tactic baseline, with authority endorsement and evidence-based arguments strongest (for gpt-4.1-nano: authority endorsement 69.39 persisted, evidence-based 68.37, logical appeal 63.27, urgency priming 66.33, anchoring 65.31, versus baseline 51.53). Logical appeal and urgency priming were comparatively weaker. In contrast, llama-3.1-8b showed high baseline susceptibility — the no-tactic condition reached 86.73 persisted and 12.76 no change, with fading near zero (0.51) — and applying tactics did not further increase persistence.

  • On-the-fly persuasion produces weak, inconsistent coding effects. Pooled coding differences between persuaded (P) and not-persuaded (NP) trials were small: TRS shifts fell in the range Δ̄ ∈ [−0.022, 0.031] with marginal evidence of separation (p ≥ 0.075), and EVS stayed near zero (|Δ̄| ≤ 0.003, p ≥ 0.553). Persona-level dispersion often exceeded the pooled means — TRS persona-level IQR was 0.064 for gpt and 0.073 for mistral, almost three times greater than each mean shift — while EVS showed smaller dispersion (IQR ≤ 0.041).

  • Aggregate web-research shifts cancel out. On-the-fly persuasion produced small, statistically insignificant changes across backbones and constructs (activity, breadth, depth). Personas within the same backbone frequently shifted in opposite directions: under gpt, persona-level breadth deltas ranged from −1.081 to +0.732 despite a pooled mean of Δ̄_Brd = −0.202 and an IQR of 0.792.

  • Long-horizon runs become longer and more active under any context injection. Relative to the no-injection baseline (C0), long-running runs showed increases in total duration (neutral +43.467, P +23.959, NP +26.849), searches (neutral +2.883, P +1.528, NP +1.926), and unique URLs (neutral +2.117, P +1.306, NP +1.327), while domain count and domain entropy changed only slightly. Persuasion-injected runs showed larger drops in query similarity (P −0.069, NP −0.089, neutral −0.028) than neutral runs.

  • Long-horizon effects are also persona-dependent. Neutral- and Claude-persona agents mostly showed positive P − NP shifts, GPT- and Qwen-persona agents mostly negative shifts, while LLaMA and Mistral were mixed.

  • Belief prefill produces the clearest behavioral shifts. Using prefill-neutral (P0) as the reference, belief-prefilled agents issued fewer searches (Δ = −1.244, 95% CI [−2.083, −0.405], p = 0.004) and visited fewer unique URLs (Δ = −0.856, 95% CI [−1.541, −1.171], p = 0.015). Tool drift increased (Δ = +1.204, CI [0.219, 2.198], p = 0.018) and the activity construct dPC_act decreased modestly (Δ ≈ −0.38, p ≈ 0.049), while dPC_brd (p = 0.231) and dPC_dpt (p = 0.461) did not change significantly. Disbelief prefill closely matched the baseline (Δ(NB − P0) = +0.076 searches, −0.032 unique URLs, +0.036 tool drift).

  • The abstract-level effect size. When belief state is explicitly specified at task time, belief-prefilled agents conduct on average 26.9% fewer searches and visit 16.9% fewer unique sources than neutral-prefilled agents — a larger and more consistent effect than on-the-fly persuasion.

  • Output quality declines are conditional, not catastrophic. Using an LLM-as-a-judge (gpt-5-mini) on a 1–5 scale, persuaded agents generally scored lower than not-persuaded agents on coverage, grounding, specificity, instruction following, and overall quality in both normal-length and long-running settings. In normal-length tasks the largest gaps appeared in coverage and specificity; in long-running tasks the clearest gap appeared in instruction following (P −0.250 versus NP −0.076 relative to no injection).

Methodology in Plain English

The authors built a four-stage pipeline — Add Persona → Persuade → Execute → Analyze — where the same agent instance both receives the persuasive content and performs the task, so any conversational state carries over. Agents are reinitialized between trials so effects do not leak across runs.

Two belief conditioning regimes. In on-the-fly persuasion, the agent is first probed for its stance on a controversial topic, then shown either nothing (C0), a neutral prompt of matched length and placement (C1), or a persuasive argument for the opposite stance (C2) written by a separate LLM writer (gpt-4.1-nano or Gemini-2.5-Flash). The agent then commits to the target stance by agreeing, restating it, and naming a concrete consideration. Stance is re-probed immediately and again at the end; successful persuasion means a change that survives. In prefilled belief conditioning, the belief (belief, disbelief, or neutrality toward a target claim) is simply stated in the message preceding the task — no probing, persuasion, or commitment — which isolates the behavioral effect of belief from persuasion mechanics.

Task execution. Each trial runs inside a multi-agent system (AutoGen) where a primary agent orchestrates decisions and supporting agents handle tool calls. Three task types are studied: an opinion change task with distractor questions sampled from WikiQA, a coding task requiring iterative debugging against test cases, and a web research task requiring open-ended search, source visits, and report synthesis. All traces — timestamps, tool use, queries, code revisions, navigation — are logged for process-level analysis.

Measurement. Coding behavior is summarized by two composite rank-based scores: the Time-and-Revision Score (TRS, aggregating coding duration, trial duration, and number of revisions) and the Edit Volatility Score (EVS, aggregating revision entropy and mean revision size, with mean-size ranks inverted to reward incremental patching). Web research behavior is summarized into three constructs — activity, breadth, and depth — via one-dimensional PCA. Long-running behavior is approximated by concatenating three web research tasks into a single continuous run, explicitly treated as a controlled proxy rather than a full simulation.

Datasets and models. Five non-control claim pairs from the Persuasion dataset with extreme human ratings for the irrelevant-persuasion topic; all 56 non-control claim pairs for the opinion change task; five gpt_difficulty=hard problems from the TACO subset of KodCode-V1; and five topics from the TREC 2014 Session Track, where agents must visit at least 5 distinct websites. Backbones are gpt-4.1-nano, mistral-nemo-12b, and llama-3.1-8b. Six persona configurations are used: Neutral, Claude, GPT, LLaMA, Mistral, and Qwen.

Why This Matters

Impact on research. The paper argues that belief-level evaluation — did the model change its stated stance? — is insufficient for agentic systems, because an expressed stance does not necessarily imply a genuine belief update that reaches decision-making. It reframes agent persuasion evaluation as a process question and shows why high variance across backbone, persona, task, and tactic means aggregated near-zero effects should not be read as absence of risk.

Real-world applications:

  • Deep research agents that synthesize reports could quietly narrow their source base after exposure to an unrelated persuasive claim, producing fluent but more brittle recommendations.
  • Coding agents that iterate on revisions could shift their debugging patterns — fewer revisions or more incremental patches — in ways invisible in the final passing code.
  • Agentic security and monitoring: because the persuasive content is task-irrelevant and may never appear in the output, output inspection alone cannot detect it, which matters when agents become targets of adversaries.
  • Long-running deployments across sessions, where small trajectory-level deviations can compound across additional searches, subtasks, or revisions.

Industry relevance. The findings point to a concrete auditing gap: deployed agents that maintain persistent context across interactions can carry belief states forward. The paper's distinction between mere agreement ("on-the-fly persuasion") and genuinely integrated belief ("belief prefill") suggests that systems which carry explicit memory or stated preferences into task context are more exposed than systems that treat each persuasion event as a transient exchange.

Future Directions

  1. Better behavioral metrics and baselines. The authors state that persuasion propagation is brittle and that future work should continue developing behavioral metrics, appropriate baselines, and distribution-aware analysis to avoid missing or misattributing agent-specific effects.

  2. From controlled proxy to real long-running deployment. The concatenated three-task run is explicitly described as an approximation, not a complete simulation, leaving open how effects compound over genuinely long horizons of 30+ tool calls across sessions.

  3. Explaining persona-level divergence. Why personas within the same backbone shift in opposite directions — and why neutral and Claude personas trend positive while GPT and Qwen trend negative — remains unexplained and is a natural target for mechanistic follow-up.

  4. Persuasion timing and integration point. A key open question is why belief specified at task time affects behavior more than belief inferred from prior conversation; the paper suggests integration into the initial context matters, but the underlying mechanism is not established.

Target Audience

This paper is most valuable to AI safety and agent-reliability researchers, LLM evaluation methodologists, and practitioners deploying tool-using agents in long-horizon settings such as research assistants or coding agents. It is also relevant to security teams concerned with context manipulation and memory poisoning, and to behavioral/social-psychology-adjacent researchers interested in how persuasion analogies transfer to machine agents. Readers should be comfortable with statistical reporting (p-values, confidence intervals, IQRs, PCA-based constructs) but do not need deep background in persuasion research.

Authors’ abstract

Modern AI agents increasingly combine conversational interaction with autonomous task execution, such as coding and web research, raising a natural question: What happens when an agent engaged in long-horizon tasks is exposed to user persuasion? Yet studying this possibility is challenging because long-running agent behavior is noisy and costly to reproduce, and it remains unclear which unique challenges emerge only in extended task execution. We study how belief-level intervention can influence downstream task behavior, a phenomenon we name persuasion propagation. We introduce a behavior-centered evaluation framework that distinguishes between persuasion applied during or prior to task execution. Across web research and coding tasks, we find that on-the-fly persuasion induces weak and inconsistent behavioral effects. In contrast, when the belief state is explicitly specified at task time, belief-prefilled agents conduct on average 26.9% fewer searches and visit 16.9% fewer unique sources than neutral-prefilled agents. These results suggest that persuasion, even in prior interaction, can affect the agent's behavior, motivating behavior-level evaluation in agentic systems.

Read the original paper