Research
Not All Answers Are Contextually Persuadable: Inference Dynamics in Large Language Models under Contextual Influence
Overview Research area: Large language model inference dynamics, contextual sensitivity, and mechanistic interpretability of prompting (with a formal/theoretical component). Technical level: Advanced.

- arXiv
- 2610.04791
- Published
- 2026-10-03
- Authors
- Zongye Hu, Weiqing Luo, Yanjie Fu, Yu Gan, Haofeng Zhang, Ziyi Huang
AI summary
Overview
Research area: Large language model inference dynamics, contextual sensitivity, and mechanistic interpretability of prompting (with a formal/theoretical component).
Technical level: Advanced. The paper develops a convergence proof for transformer inference under repeated context (Cesàro limits, RoPE logit structure, layer-wise induction) alongside an empirical validation suite spanning six models and seven benchmark subsets.
Scope: A theoretical and empirical study of what happens to an LLM's internal inference trajectory — and its final answer — when an identical contextual assertion is repeated an unbounded number of times in a single prompt.
What This Paper Is About
Prompting, in-context learning, and answer-steering methods all rely on the assumption that LLMs adjust their predictions to inference-time context. A widespread extension of that assumption is that repeated contextual assertions behave like accumulating evidence: repeat a candidate answer enough times and the model should eventually be pushed toward it. This paper asks what the asymptotic behavior of LLM inference actually is under unbounded contextual repetition, characterizing the internal representation-level dynamics rather than only the observable answer flips.
Key Contributions
-
Problem formulation. The authors recast contextual influence as a problem of internal inference dynamics, introducing a formal framework that models how repeated contextual signals shape representation-level inference trajectories, moving beyond answer-level prediction changes to a mechanistic account.
-
Asymptotic inference convergence. In a controlled single-round setting (a fixed query followed by repeated copies of an identical assertion, with no additional evidence, reasoning steps, or multi-turn interaction), they establish formal convergence guarantees showing that internal inference trajectories converge to stable, query-dependent limits rather than drifting without bound.
-
Representation–prediction alignment. The theoretical predictions are validated empirically across models and tasks, showing a tight correspondence between representation-level dynamics and observable prediction behavior, and identifying when prediction changes become inevitable versus provably unattainable under unbounded repetition.
-
A practical estimator. They derive an answer-shift metric in latent space plus a Monte Carlo estimator that approximates the infinite-repetition limit from a single forward pass at a large finite repetition count, decomposing the limit into per-layer attention and FFN contributions.
Main Findings
-
Inference does not drift without bound. Under the stated assumptions, the target representation at the queried position converges to a well-defined finite limit at every layer as the repetition count N grows (Theorem 3.6). Because next-token logits follow from the final-layer hidden state, the next-token probability distribution converges as well.
-
Repetition is not accumulating evidence. Repeated assertions do not act as progressively accumulating evidence. Their influence is bounded by the structure of the model's internal representations, so a reiteration-based nudge can fail even under unbounded repetition.
-
Outcomes are heterogeneous across questions and models. Figure 1 illustrates qualitatively distinct behaviors under repeated contextual assertions: immediate answer flips, delayed changes, early saturation, and complete invariance to repetition. Some predictions flip inevitably; others are provably unattainable.
-
The convergence mechanism is RoPE plus the repeated template. The proof rests on the observation that RoPE attention logits at a fixed layer and head form a finite trigonometric polynomial (equivalently, a finite sum of complex exponentials) in the relative distance, so the softmax normalizer and the value-weighted numerator have well-defined Cesàro limits. At the base layer the loop key/value vectors are exactly periodic with period T, and the base-layer output at loop positions becomes asymptotically nearly periodic, which lets the argument be propagated layer by layer.
-
A single structural property drives convergence. The convergence argument relies on the target position being preceded by infinitely many copies of the loop, which allows averaging over relative positions to stabilize.
-
Answer preference is a logit gap. The final-layer answer-shift metric — the inner product of the contrast direction between the asserted answer and the reference answer with the residual stream — is exactly the logit gap between the two answers, and its sign determines whether unbounded repetition can flip the prediction.
-
The Monte Carlo estimator tracks forward computation. Figure 2 compares representation-level predicted answer shifts against forward-computed output-level shifts across multiple models on OpenBookQA at N = 1000; the points line up along the diagonal, indicating high quantitative fidelity.
-
Predictive distributions stabilize. Figure 3 shows KL divergence trajectories of next-token predictive distributions under repetition; across models the KL divergence stabilizes as repetition length increases.
-
Layer-wise convergence rates at N = 1000. Table 1 reports layer-wise convergence of inference dynamics at N = 1000 on OpenBookQA, MINTAKA, and SimpleQA at divergence levels 0.1, 0.05, and 0.01, separately for attention and FFN contributions. Convergence percentages are generally in the high 90s at the 0.1 level and decline at the stricter 0.01 level; the lowest value shown is 33.9 (FFN) for Qwen2.5-1.5B on OpenBookQA at 0.01, with its attention value at 47.6.
-
Predicting preference changes. Table 2 evaluates prediction of answer preference changes in the Transfer, Correct, and Mislead regimes at N ∈ {500, 750, 1000}, reporting accuracy and F1. In the table's mean row, Transfer is 85.5/91.5, 85.7/91.8, and 86.1/92.0 (ACC/F1 at N = 500, 750, 1000), Mislead is 82.6/86.1, 80.4/84.6, and 80.0/84.6, and Correct is 80.8/78.9, 79.9/78.9, and 82.2/89.8.
-
Three practical implications. Inference dynamics: repetition only matters when it aligns with malleable inference directions, otherwise inference converges to a regime where further repetition is ineffective. Robustness and susceptibility: susceptibility is not uniform but depends systematically on both the query and the model, so robustness assessments should separate intrinsically stable inputs from vulnerable ones. Repetition versus evidence: repeated assertions alone do not constitute accumulating evidence during inference.
Methodology in Plain English
The authors isolate repetition as the only source of contextual signal. A prompt is split into three parts: a fixed query/prefix (y), a fixed loop template (x) repeated verbatim N times, and a suffix (z) containing the answer cue. The target position is chosen so it always indexes the same suffix token regardless of N, even though its absolute position grows linearly with N.
They then analyze a standard decoder-only transformer with pre-norm attention and feed-forward blocks, Rotary Position Embedding (RoPE), and an assumed infinite context window, under a set of regularity assumptions (deterministic evaluation, causality, no absolute-position term outside RoPE, finite operator norms, uniformly bounded residual streams, and Lipschitz layer-norm and feed-forward blocks).
The proof proceeds in five steps: show RoPE logits are finite trigonometric polynomials in the relative distance; use the base-layer periodicity of loop key/value vectors to establish Cesàro limits for each attention head's output; lift head-wise limits to the full base-layer residual stream using multi-head aggregation and continuity of the normalization and feed-forward blocks; show the base-layer output at loop positions is asymptotically periodic so the periodic loop structure survives into the next layer; then close a layer-wise induction up to the final layer via a perturbation–reduction argument.
To connect the math to behavior, they define an answer-shift metric: the projection of the residual stream onto the difference between the output-head rows of the asserted answer and the reference answer. At the final layer this projection equals the logit gap between the two answers, and by the convergence theorem it has a limit whose sign tells whether repetition can ever flip the prediction.
Because the limit is not directly computable, they approximate it from a single forward pass at a large finite N. The RoPE logit depends on a past position only through relative distance, so they treat the last loop block as a template and model the loop contribution as an integral over the RoPE phase angle, estimated by Monte Carlo sampling of copy offsets and template indices with rescaling by NT/D (unbiased, and exact once D ≥ NT). Non-loop positions are added exactly. The resulting residual stream telescopes into per-layer attention and FFN terms, giving an additive decomposition of the predicted answer shift.
Experiments cover seven benchmark subsets — OpenBookQA (closed, 1,769 QA), MINTAKA (open, 370 QA), SimpleQA (open, 191 QA), plus subsets of SYCON-Bench, Farm, BeHonest, and sycophancy-eval — and six LLMs: Falcon3-7B-Base, Mistral-7B-v0.1, Apollo-1-4B, Qwen3-4B, Qwen2.5-1.5B, and Falcon3-3B-Base, spanning the Qwen, Mistral, and TII families and large (>6B), medium (3B–6B), and small (<3B) parameter scales. The paper notes that the provided content is truncated at Section 5.2, so the remaining experimental results are not reported here.
Why This Matters
Impact on research: The paper challenges a working assumption baked into empirical prompting practice and evaluation protocols — that sufficient repetition will always eventually change a model's answer. It replaces that assumption with a measurable, query- and model-dependent limit, and it links asymptotic analysis of transformer computation (previously largely task-agnostic) to concrete prediction behavior on real question-answering tasks.
Real-world applications:
- Factual question answering and evaluation benchmarks: Knowing which inputs are intrinsically stable versus which can be flipped by mere repetition informs how benchmark results and evaluation protocols should be interpreted.
- Adversarial prompting and prompt injection: Repeat-based steering and spam-style injection attempts have a computable ceiling; the answer-shift metric and sign test can flag inputs that are provably safe from this class of manipulation and those that are not.
- Steering and prompting method design: Answer priming, instruction reinforcement, and repetition-based prompting can be designed and budgeted using the estimate of how much shift the unbounded-repetition limit actually permits.
- Safety-critical deployments: The framework lets developers distinguish between contexts where repeated assertions are ineffective and contexts where a prediction change is inevitable, supporting more targeted guardrails than blanket robustness claims.
Industry relevance: The work connects to model development practice — the authors explicitly frame their findings as having practical implications for model development, and robustness assessments that separate intrinsically stable inputs from vulnerable ones. The author list includes an industrial affiliation (Morgan Stanley), and the conclusions are relevant to anyone deploying LLMs where spurious or adversarial contextual cues should not change answers.
Future Directions
- Beyond single-round, single-assertion repetition. The analyzed setting deliberately excludes additional evidence, reasoning steps, and multi-turn interaction. Extending the framework to dialogue dynamics and sustained interactional pressure — which the related work attributes to training incentives and dialogue dynamics rather than inference-time repetition — is a natural next step.
- Relaxing the theoretical assumptions. The results rest on Assumption 3.1, including RoPE as the only positional mechanism, an infinite context window, deterministic evaluation, bounded residual streams, and Lipschitz layer-norm and feed-forward blocks. Behavior under other positional schemes, sampled decoding, or unbounded residual streams is not established.
- Characterizing what makes an answer contextually persuadable. The paper shows malleability varies systematically across queries and models but does not derive a general predictor of which directions are malleable; identifying that structure would directly serve robustness assessment.
- Validating the finite-N approximation at larger budgets. The infinite-repetition limit is estimated from a single forward pass at finite N (with experiments at N = 500, 750, and 1000, and estimator tables at N = 1000), and the Monte Carlo estimator is exact only when D ≥ NT. How well the approximation holds as N and D grow, and how far it extends across the broader benchmark and model suite, is not fully reported in the available content.
Target Audience
Researchers and graduate students in machine learning working on LLM interpretability, prompting, in-context learning, and robustness/adversarial behavior; theoretically inclined readers comfortable with transformer architecture details, RoPE, and asymptotic limit arguments; and practitioners in model evaluation, red-teaming, and safety — particularly those who design or defend against answer-steering and repeat-based prompting techniques.
Authors’ abstract
At the core of modern prompting techniques is contextual sensitivity, the ability of large language models to adapt their predictions based on inference-time context. Despite its central role, inference behavior under strong contextual influence remains poorly understood, particularly at the level of internal inference dynamics. We introduce a theoretical framework for analyzing contextual influence through inference dynamics, enabling quantitative characterization of inference behavior beyond output-level answer changes. Our analysis shows that inference dynamics do not exhibit unbounded drift under repeated contextual assertions. Instead, predictive representations converge to stable, query-dependent regimes that fundamentally constrain whether contextual signals can alter a model's prediction. This leads to a surprising finding: Repeated contextual assertions do not act as accumulating evidence during inference and may therefore fail to alter a model's prediction even under unbounded repetition, while in other cases a prediction change becomes inevitable. We empirically validate our theoretical predictions, demonstrating strong alignment between theory and observed inference behavior. These contributions offer a principled pathway toward characterizing the limits of contextual influence during inference, providing practical implications for model development.