Research
Depth-Wise Activation Steering for Honest Language Models
Overview Research area: Machine learning — mechanistic interpretability and activation steering (representation engineering) for AI honesty and safety alignment. Technical level: Intermediate. The pap

- arXiv
- 2512.07667
- Published
- 2025-12-08
- Authors
- Gracjan Góral, Marysia Winkels, Steven Basart
AI summary
Overview
Research area: Machine learning — mechanistic interpretability and activation steering (representation engineering) for AI honesty and safety alignment.
Technical level: Intermediate. The paper assumes familiarity with transformer residual streams, inference-time intervention, PCA, and LoRA, but the core idea (a smooth schedule over layers) is explained in accessible terms.
Scope: The paper proposes a training-free method that distributes a fixed activation-steering budget across network depth with a Gaussian curve, and evaluates whether it improves honest reporting on the MASK benchmark across seven open-weight models.
What This Paper Is About
Large language models sometimes state things they internally know to be false. That is an honesty failure, not a knowledge failure, and it undermines oversight because a model that knows but misreports is harder to audit. The paper's goal is a cheap, test-time control knob that steers a model toward reporting what it already represents, without fine-tuning, by choosing where in the network depth the intervention is applied.
Key Contributions
- A depth-allocation framing. The paper formulates honesty-directed activation steering as a problem of allocating a fixed intervention budget across network depth, rather than picking one layer or spreading strength uniformly.
- An analytic Gaussian depth schedule. It introduces a simple, training-free schedule with per-layer strength
α_ℓ = exp(-(ℓ - μ)² / (2σ²)), parameterized by center μ and width σ, requiring no finetuning. - Measured honesty gains across model families. On MASK, Gaussian scheduling improves honesty over both no-steering and single-layer baselines in six of seven open-weight models spanning LLaMA, Qwen, and Mistral families.
- Equal-budget ablations isolating distribution shape. Holding total steering norm fixed on LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, the Gaussian schedule outperforms random, uniform, and box-filter depth allocations.
- Evidence of complementarity with parameter-efficient fine-tuning. After LoRA fine-tuning, scheduled activation control remains competitive, suggesting it complements rather than merely substitutes for such training.
Main Findings
- Routine gains instead of occasional ones. Across seven open-weight models, the Gaussian schedule improves honesty over both no-steering and single-layer baselines in six of seven cases. The authors state the conclusion section reports superiority to single-layer baselines in five of seven models — a discrepancy with the six-of-seven figure given in the abstract and results.
- Double-digit absolute gains. LLaMA-3.1-8B-Instruct rises from 20.8 to 38.0 (+17.2) honesty relative to no-steering, and Mistral-7B-Instruct-v0.2 rises from 18.9 to 32.9 (+14.0).
- Single-layer edits can backfire, and the schedule reverses the damage. On Qwen-2.5-7B-Instruct, single-layer steering lowers honesty below no-steering (27.0 → 24.4), and on Qwen-2.5-14B-Instruct it does the same (23.4 → 20.9). The Gaussian schedule lifts these to 33.9 and 30.1 respectively.
- One clear counterexample. On Mistral-7B-Instruct-v0.2, single-layer steering reaches a higher peak than Gaussian (40.7 vs. 32.9). Even there, the Gaussian schedule still beats no-steering by +14.0.
- Depth distribution shape matters independently of total strength. Under equal-budget allocations, the Gaussian schedule achieves the highest MASK honesty gains on LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, beating random, uniform, and box-filter distributions.
- Scale trend within LLaMA. For LLaMA models, the scheduler increases honesty consistently as model size grows.
- Competitive with LoRRA fine-tuning. A LoRRA-style LoRA fine-tune that internalizes the same honest-vs.-dishonest targets improves honesty over no-steering on both LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, but the Gaussian schedule still delivers the largest gains.
- Taken together, four takeaways: a single training-free scheduler yields consistent improvements across families; it prevents degradations single-layer edits induce; where you steer matters under a fixed budget; and benefits persist alongside LoRA.
Methodology in Plain English
The researchers first build a direction per layer that points toward "honest" behavior. For each transformer block (excluding the embedding layer), they take pairs of prompts designed to elicit honest versus dishonest responses, record the residual-stream activations at the last non-padding token, subtract the dishonest activation from the honest one, stack those differences into a matrix ΔA_ℓ ∈ ℝ^{n×d}, and apply one-component PCA. The first principal axis becomes that layer's steering direction d_ℓ, oriented so that it aligns positively with held-out difference vectors.
At inference, instead of adding that direction at one layer or at equal strength everywhere, they add a scaled residual δ_ℓ = α_ℓ d_ℓ at each block, where α_ℓ follows a normalized Gaussian centered at μ = ⌊L/2⌋ with width σ > 0. Strength is weak in early layers, peaks in mid-to-late layers where they argue abstract semantic features are better separated, and tapers near the output.
They evaluate on MASK, which pairs each factual proposition with a ground-truth label, an adversarial pressure prompt that incentivizes a false answer, and a neutral belief-elicitation prompt. Seven open-weight models are tested: Llama 3.2 (1B and 3B)-Instruct, Llama 3.1 8B-Instruct, Qwen 2.5 (3B, 7B, and 14B)-Instruct, and Mistral-7B-Instruct-v0.2.
Baselines are vanilla inference with no steering and single-layer steering. Hyperparameters are chosen by grid search — over intervention layer and steering coefficient for single-layer steering, and over peak μ and standard deviation σ for the Gaussian method — using 25% of MASK as a validation split. Outputs are mapped to the benchmark's discrete label space with gpt-oss-20B at temperature 1.0, using the judging prompts from Ren et al. (2025). Honesty is computed by eliciting both a pressured statement and a neutral belief, mapping both to proposition values, and comparing the statement against the belief, averaged over the full benchmark.
Why This Matters
Impact on research. The paper reframes activation steering as a budget-allocation problem over depth rather than a search for a single "best layer." Its equal-budget ablations give direct evidence that the shape of the depth distribution matters beyond total intervention strength, which is a reusable design insight for other steering targets such as sycophancy, toxicity, hallucination, or refusal. It also explicitly targets honesty as distinct from factual accuracy, aligning with benchmarks like MASK that separate the two.
Real-world applications:
- Test-time honesty controls for open-weight models deployed without a retraining pipeline.
- Auditing and oversight tooling, where a model that "knows but misreports" can otherwise evade detection.
- A complement to existing safety stacks such as RLHF, Constitutional AI, and external safety classifiers, rather than a replacement.
- A low-cost fallback where a full fine-tune is infeasible but the targets used to build control vectors are already available.
Industry relevance. The method is model-agnostic, requires no finetuning, and is described as a low-cost control knob, which makes it attractive for organizations that serve open-weight models and need an inexpensive, deployable intervention. The authors release code and experiments at https://github.com/marysia/gaussian-activation-steering. The requirement for per-layer activation access means it applies to open-weight models rather than closed API-only systems.
Future Directions
- Broaden benchmark coverage. The authors note their experiments focus primarily on MASK and call for validation across a broader range of safety benchmarks, naming Machiavelli (Pan et al., 2023) as an example.
- Strengthen evaluation robustness. Because the study relies on an external LLM judge with specific prompting strategies, the authors recommend incorporating multiple judges or human evaluation, since results may be sensitive to judge selection and prompt design.
- Resolve the single-layer counterexample. Mistral-7B-Instruct-v0.2 is the one model where single-layer steering peaks higher than Gaussian (40.7 vs. 32.9); understanding when a point edit beats a smooth schedule remains open.
- Determine when schedules help versus where fine-tuning is required. The paper shows scheduled steering remains competitive alongside LoRRA fine-tuning on LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, but the division of labor between the two is not settled.
- Extend beyond open-weight settings. The method's need for activation-level access restricts its applicability, leaving open how honesty steering could be offered for models without such access.
Target Audience
Alignment and safety researchers working on mechanistic interpretability, activation steering, or representation engineering; practitioners deploying open-weight models who need inference-time honesty controls without retraining; and engineers building evaluation or oversight pipelines around benchmarks like MASK. Readers will benefit most if they are comfortable with transformer internals, residual streams, and the distinction between a model's knowledge and its willingness to report it.
Authors’ abstract
Large language models sometimes assert falsehoods despite internally representing the correct answer, failures of honesty rather than accuracy, which undermines auditability and safety. Existing approaches largely optimize factual correctness or depend on retraining and brittle single-layer edits, offering limited leverage over truthful reporting. We present a training-free activation steering method that weights steering strength across network depth using a Gaussian schedule. On the MASK benchmark, which separates honesty from knowledge, we evaluate seven models spanning the LLaMA, Qwen, and Mistral families and find that Gaussian scheduling improves honesty over no-steering and single-layer baselines in six of seven models. Equal-budget ablations on LLaMA-3.1-8B-Instruct and Qwen-2.5-7B-Instruct show the Gaussian schedule outperforms random, uniform, and box-filter depth allocations, indicating that how intervention is distributed across depth materially affects outcomes beyond total strength. The method is simple, model-agnostic, requires no finetuning, and provides a low-cost control knob for eliciting truthful reporting from models' existing capabilities.