Research
Towards Understanding Steering Strength
Overview Research area: Post-training control of large language models, specifically activation steering and its theory. The work sits at the intersection of mechanistic interpretability, the linear r

- arXiv
- 2602.02712
- Published
- 2026-02-02
- Authors
- Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
AI summary
Overview
- Research area: Post-training control of large language models, specifically activation steering and its theory. The work sits at the intersection of mechanistic interpretability, the linear representation hypothesis, and learning theory.
- Technical level: Intermediate. The paper is a theory paper with proofs in appendices, but its questions and headline results can be understood with basic familiarity with language model internals (residual streams, next-token probabilities, cross-entropy) and the steering recipe of adding a vector to a hidden representation.
- One-sentence scope: The paper gives the first theoretical analysis of the steering strength α, characterizing how it shapes next-token probabilities, the probability of a concept in the output, and cross-entropy, then validates these predictions on eleven language models ranging from a small GPT architecture to modern models.
What This Paper Is About
Activation steering is a popular way to control a language model after training: pick a direction in the model's internal activation space that corresponds to a concept, and at inference time shift the residual stream by α times that direction, h ← h + α v. While many methods exist for choosing the direction v, the magnitude α is poorly understood even though its importance is obvious — too little and the intended behavior never appears, too much and the model's performance degrades beyond repair.
This paper builds a theoretical account of α. Using a simplified but analytically tractable model of a transformer trained on a synthetic concept-structured dataset, it derives precise qualitative laws for how α controls next-token probabilities, concept presence, and cross-entropy, and then tests those laws on real language models.
Key Contributions
- Characterization of α's effect on three quantities: the paper proves how steering strength affects next-token probabilities (Theorem 3.3), the probability of a concept in the model's output (Theorem 3.6), and cross-entropy (Theorem 3.8).
- Formalization of the practical steering setup and a large-α limit: it specifies a real-life activation steering setup for decoder-only transformers (attention, feed-forward, and normalization blocks) and derives the large-α limit of next-token probabilities for a transformer (Proposition 4.1).
- Empirical validation across modern models: the theoretical predictions are validated empirically on eleven language models, ranging from a small GPT architecture to modern models (Section 5).
- Discovery of non-monotonic behavior: the analysis reveals surprising effects, in particular that most token probabilities rise, peak, and then fall as α grows, rather than increasing monotonically with steering.
The authors report that the code for all experiments is available at https://github.com/MagamedT/steering.
Main Findings
- Most tokens exhibit a "bump" (Theorem 3.3): for any token z that is neither a maximum-log-odds nor a minimum-log-odds token, there is a unique α_(j,z) such that the probability increase Δp(z | c_j, α) is strictly increasing on (−∞, α_(j,z)] and strictly decreasing on [α_(j,z), +∞). Probability rises, peaks, then falls as α grows.
- Off-target tokens peak earlier than target tokens (Theorem 3.3): for any target token z ∈ T and off-target token z′ ∉ T, the peak satisfies α_(j,z′) < α_(j,z). As α increases, off-target probabilities start to fade while target probabilities are still rising, which helps steering stay focused on the target concept. This also implies a steering "sweet spot" where target tokens are favored but the distribution has not collapsed onto a few tokens.
- A few tokens are exceptions (Theorem 3.3): tokens attaining the maximum log-odds (and belonging to the target concept) increase strictly on all of ℝ; tokens attaining the minimum log-odds (and outside the target concept) decrease strictly.
- The bump location depends on the context: α_(j,z) varies across contexts c_j, which the authors say suggests α should be chosen adaptively with respect to the input prompt, as proposed in prior adaptive-steering work (Hedström et al., 2025; Ferrando et al., 2025).
- Concept probability follows a sigmoidal / tanh shape (Theorem 3.6): the concept probability increase takes the form Δp(C | α) = (1/(2|C|))(tanh((ν_j(α) + r_j)/2) − r′j). The target concept Δp(T | α) is increasing in α. For any concept C′ ≠ T containing neither maximal nor minimal log-odds tokens, Δp(C′ | α) converges to −(1/|C′|) Σ{z ∈ C′} p(z | c_j) as α → ±∞. For any concept C ≠ T whose tokens all have log-odds below those of the remaining tokens, Δp(C | α) is decreasing in α.
- This matches a known empirical trend but partly disagrees with prior theory: the sigmoidal result is consistent with the tanh(α) trend observed empirically by Von Rütte et al. (2024). It slightly disagrees with Park et al. (2024, Theorem 2.5), which predicts target-concept probability increases while off-target concept probability stays constant. The authors attribute the difference to model assumptions and to their fine-grained, token-level definition of concept increase, and note that sampling-based concept metrics can mask changes that occur among low-probability tokens.
- Cross-entropy is locally U-shaped and steering necessarily degrades performance (Theorem 3.8): as α → 0, ΔCE(α) = (1/2) Σ_{j ∈ [m]} π_j Var_j(M(Z)) α² + o(α²). There is no linear term and the quadratic coefficient is a variance of log-odds, hence nonnegative. The authors describe this as the first theoretical characterization of how a performance measure varies with steering strength, and as a theoretical justification for the prior empirical observation that ΔCE(α) is locally quadratic in α.
- Large-α behavior: in the theoretical setting, Δp(α) concentrates on the tokens attaining maximum log-odds as α → +∞ and on the minimum-log-odds tokens as α → −∞ (Proposition B.1). For modern LLMs the limits are characterized in Proposition 4.1, and the cross-entropy curves are reported to plateau for large |α|.
- Positive steering drives the bump for target tokens: with the dataset used in the experiments (Appendix B.3), the bump pattern for target tokens occurs only for positive α, matching the intuition that positive steering increases their probabilities (Remark 3.4).
- Empirical confirmation on real models: for the concept "evil," the next-token probability curves of the eight highest-probability tokens at α = 200 show the predicted pattern — most tokens bump, a few increase throughout — and the selected tokens are all related to the steered concept (Figure 6). Across models, concept probability for the concepts depression, joy, and evil, estimated with a judge LLM (Gemma 3 12B), shows the predicted sigmoidal trend, and cross-entropy is locally U-shaped around α = 0 (Figure 7). The specific models among
Authors’ abstract
A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations. Namely, identify a well-chosen direction depending on the task at hand and perturbs representations along this direction at inference time. While many propositions exist to pick this direction, considerably less is understood about how to choose the magnitude of the move, whereas its importance is clear: too little and the intended behavior does not emerge, too much and the model's performance degrades beyond repair. In this work, we propose the first theoretical analysis of steering strength. We characterize its effect on next token probability, presence of a concept, and cross-entropy, deriving precise qualitative laws governing these quantities. Our analysis reveals surprising behaviors, including non-monotonic effects of steering strength. We validate our theoretical predictions empirically on eleven language models, ranging from a small GPT architecture to modern models.