Skip to content
AI.info

Research

There Is More to Refusal in Large Language Models than a Single Direction

Overview Research area: Mechanistic interpretability and alignment of large language models (LLMs), specifically how refusal and other non-compliant behaviors are represented inside a model's residual

There Is More to Refusal in Large Language Models than a Single Direction
arXiv
2602.02132
Published
2026-02-02
Authors
Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, Husrev Taha Sencar

AI summary

Overview

Research area: Mechanistic interpretability and alignment of large language models (LLMs), specifically how refusal and other non-compliant behaviors are represented inside a model's residual stream (listed on arXiv under Natural Language Processing).

Technical level: Advanced. The paper assumes familiarity with residual-stream activations, activation steering, directional ablation, and sparse autoencoders (SAEs), although its core argument can be followed with a conceptual understanding of these tools.

Scope: The paper tests whether the widely used "single refusal direction" account of LLM refusal holds up across 11 categories of refusal and non-compliance in three instruction-tuned models, and uses sparse autoencoders to explain what linear interventions hide.

What This Paper Is About

Prior work (Arditi et al., 2024) argued that refusal in LLMs is mediated by a single linear direction in activation space, which can be amplified to induce refusal or removed ("abliteration") to suppress it. This paper shows that account is incomplete: different refusal and non-compliance categories correspond to geometrically distinct directions, yet steering along any of them produces nearly identical refusal and over-refusal trade-offs. The goal is to reconcile the apparent simplicity of refusal steering with the observed diversity of refusal behaviors, and to characterize the internal feature structure that linear control flattens.

Key Contributions

  1. A broader reframing of refusal. The authors treat refusal as part of a wider class of non-compliant behaviors, covering 11 evaluation splits across four benchmarks (WildGuardMix, SorryBench, CoCoNot, XSTest) spanning four regimes: safety and content-policy, capability limitation, underspecification, and over-refusal.
  2. Evidence that refusal directions are geometrically distinct but category-structured. Pairwise cosine similarities between the 11 directions typically fall in the 0.4–0.6 range, with several pairs close to orthogonal, and hierarchical agglomerative clustering groups the directions along the refusal categories.
  3. Demonstration of a shared one-dimensional control knob. Steering along any of the 11 directions produces broadly similar refusal-rate and over-refusal-rate increases, so different directions change how the model refuses rather than whether it refuses.
  4. An SAE-level explanation. Sparse autoencoders reveal a reusable core of refusal latents shared across splits plus a long tail of style- and domain-specific latents, and SAE-derived refusal directions reconstruct the activation-space directions with high cosine similarity.

Main Findings

  • Distinct directions, similar behavior. On Gemma-2-9B-IT, refusal directions learned from the 11 splits have typical pairwise cosine similarities of 0.4–0.6, yet steering along each direction on a controlled test set of 200 prompts (50 each of HR, HC, BR, BC) pushes refusal rates on harmful prompts and over-refusal rates on benign prompts up together toward saturation.
  • The trade-off is remarkably uniform. Steering results are reported for Gemma at α = 100, Llama at α = 2.5, and Qwen at α = 15. For example, under Gemma steering, Humanizing-CCN reaches accuracy 0.515, refusal rate 1.000, over-refusal rate 0.970; the CoCoNot non-safety categories saturate slightly more gradually than SorryBench- and WildGuard-derived directions, but follow the same trade-off. The lowest refusal rate in the steering table is Llama on Incomplete-CCN (refusal rate 0.650, over-refusal rate 0.700).
  • Ablation tells a more nuanced story. Ablating the safety-derived direction suppresses safety-style refusals (over-refusal on benign prompts drops toward zero, and refusal rates on SorryBench- and WildGuard-style splits collapse), but CoCoNot non-safety splits retain residual refusal of roughly 0.3–0.5, suggesting additional refusal-relevant components.
  • A mixed-source ablation direction generalizes broadly. When the ablation direction is built from a round-robin pool sampled across all 11 splits, ablating it yields 87.1% (Qwen), 69.5% (Llama3) and 81.1% (Gemma2) of baseline-refused harmful prompts flipped to non-refusal, compared with much weaker transfer for individual directions (for instance, Humanizing-CCN: 1.3 / 19.5 / 0.4 and SafetyCore-WGM: 12.5 / 14.1 / 1.8 across Qwen / Llama3 / Gemma2).
  • Refusal style varies systematically. Despite similar rates, the form of refusal differs by direction: humanizing directions emphasize the model's non-human nature, incomplete-request directions produce terse confusion, indeterminate or unsupported directions stress impossibility or lack of agency, SorryBench directions foreground moral judgments or professional disclaimers, and WildGuard/XSTest directions emphasize safety, illegality, or harm.
  • SAE latents recover the same knob. SAE-based refusal steering (top K = 10 latents averaged into a direction) monotonically raises both refusal and over-refusal, with selected β = 60 for Gemma and β = 1.2 for Llama. Random-latent controls produce much weaker, non-monotonic changes.
  • SAE directions align with activation-space directions. Cosine similarity between the SAE-induced refusal direction and the corresponding activation-space direction is uniformly high, typically 0.85–0.97 across splits (per-split means range from 0.849 with σ = 0.209 for Safety-CCN to 0.971 with σ = 0.039 for Humanizing-CCN).
  • A shared core plus a long tail. Latent overlap analysis across the 11 splits reveals a small reusable core of shared refusal latents plus a longer tail appearing in only one or a few splits, aligning with specific refusal styles (such as terse "I don't understand" replies) or domains (such as legal disclaimers in unqualified advice).
  • Some latents are polysemous. LLM-assisted annotation of top-10 recurring latents surfaces domain-general features (harmful requests, prohibited or illicit content, disinformation, policy-violating directives), while others are subcategory-specific; several shared latents receive harm-focused interpretations in SorryBench but capability-focused interpretations in CoCoNot.
  • The unsteered baseline. On the balanced four-pool controlled test set, the unsteered base model attains 50% on accuracy, refusal rate and over-refusal rate by construction, since correct outcomes are confined to HR and BC.

Methodology in Plain English

The authors work at two levels of the model's internal representation.

Level 1 — activation space. For each evaluation split, they collect residual-stream activations at a fixed token position immediately before the assistant's response, average them separately over non-compliant prompts and benign prompts, and define the refusal direction as the difference between the two centroids. They then intervene in two ways: induction (adding the direction, scaled by a strength α) and ablation (projecting the direction out of the residual stream). The ablation is applied at every layer, while steering is applied at the selected layer. Steering strength is chosen by grid search over a small set such as {5, 10, 20, 30, 60}, picking the smallest value that reaches at least a 90% refusal rate on harmful prompts while keeping benign over-refusals below a threshold.

Level 2 — sparse autoencoder feature space. Pretrained JumpReLU residual-stream SAEs decompose activations into sparse latents with associated decoder directions. The authors hook the residual stream at the token immediately preceding the assistant's response, which they call the model's "decision state." To find refusal latents, they compare each latent's firing rate on harmful prompts the base model correctly refused (HR) with its firing rate on benign prompts the base model correctly complied with (BC), using the difference as a "refusal separation score," and take the top-K positive-scoring latents per split and layer. Averaging the decoder directions of those latents gives an SAE-based refusal direction, which they use for steering; they also ablate by zeroing selected latents and decoding. Controls are built from random latent subsets of the same size K and from random unit vectors. They select K = 10 based on a validation sweep over K ∈ {1, …, 15}, where performance peaks near 10.

Models and data. Experiments use gemma-2-9b-it (activations read at layer 20, position −2, steering at the same layer with α = 100), Meta-Llama-3-8B-Instruct (steering at layer 16, position −2, α = 2.5; ablation at layer 12, position −5, the eot_id token) and Qwen-7B-Chat (ablation and steering at layer 14, position −1). Directions are estimated from 32 harmful and 32 benign prompts per split. SAE analysis uses GemmaScope SAEs at layers ℓ ∈ {9, 20, 31} for Gemma and a Llama-3.1-specific residual-stream SAE release via SAELens at a fixed late layer; because no public SAE matches the exact Qwen-7B-Chat checkpoint used, SAE analysis is restricted to Gemma and Llama. Outputs are labeled as refusals or compliances with the WildGuard refusal classifier, with manual validation, and evaluated by overall accuracy, refusal rate and over-refusal rate.

Why This Matters

Impact on research. The paper complicates a now-standard story in interpretability: that a single linear direction governs refusal and that removing it is a complete account of abliteration. It argues that linear interventions act as a coarse behavioral knob that collapses a structured internal system — a shared refusal core plus style- and domain-specific features — into one effective dimension. That is a pointed limitation of linear interpretability for alignment-relevant behavior, and it connects to prior work on multiple refusal directions (Zhang and Sun, 2026; Wollschläger et al., 2025), dormant "hydra" features (Prakash et al., 2026) and separated harm/refusal feature sets (Yeo et al., 2025).

Real-world applications (as motivated by the paper's framing):

  • Safety evaluation and red-teaming of aligned models. Understanding that refusal directions are category-structured helps auditors test whether safety-style and non-safety refusals are being suppressed independently.
  • Diagnosing over-refusal in deployed assistants. The XSTest split and the over-refusal metric directly target benign prompts that are wrongly declined, which is a common user-facing complaint.
  • Designing non-compliance behavior. Because direction choice changes refusal style (clarification, capability disclaimers, moral framing), developers can think about which refusal form is appropriate for which request type.
  • Assessing the robustness of alignment mitigations. The results underscore that steering-based interventions are not suited as principled safety mechanisms.

Industry relevance. Abliteration-style attacks are widely used to strip refusal behavior from open-weight models. This work shows that removing a safety direction leaves a residual refusal component from other categories (roughly 0.3–0.5 on CoCoNot non-safety splits), but also that a mixed-source direction derived from all 11 splits suppresses refusal broadly (87.1% / 69.5% / 81.1% flipped for Qwen, Llama3 and Gemma2). That distinction matters for anyone relying on refusal robustness as a safety property.

Future Directions

  • Extending to earlier layers and pre-commitment positions. The authors hook only the decision-state token; they note that recent work (Zhao et al., 2025) suggests safety- and harmlessness-related directions may already be encoded at earlier layers and at the final user-input token, and that extending the shared-core analysis there would disentangle input-driven from output-driven latents.
  • Testing larger and more diverse models. The authors explicitly make no claim about substantially larger models (e.g., 70B+), non-decoder architectures, or base models without instruction tuning, where the refusal mechanism itself is much weaker (Kissane et al., 2024).
  • Broadening SAE coverage. Sparse-feature analysis is limited by the availability of publicly released SAEs, which restricts the analysis to a limited set of layers and models and may bias it toward later layers where SAE coverage is more common.
  • Moving from linear control to feature-level alignment. The paper concludes by motivating more robust, feature-level approaches to alignment rather than relying on steering directions, which its results characterize as a coarse and brittle knob.

Target Audience

Interpretability and alignment researchers, particularly those working on activation steering, directional ablation/abliteration, and sparse-autoencoder analysis of safety behavior. It is also relevant to safety engineers and policy-oriented practitioners who need to reason about how robust refusal behavior is in open-weight instruction-tuned models, and to graduate-level students who already have some grounding in transformer internals. Readers looking for a first introduction to mechanistic interpretability would find the paper's notation and interventions demanding.

Code is available at https://github.com/fjoad/more-than-one-direction.

Authors’ abstract

Prior work argues that refusal in large language models is mediated by a single direction, enabling steering and abliteration. We show that this account is incomplete: across diverse refusal and non-compliance categories, refusal behaviors correspond to geometrically distinct directions in activation space. Yet activation steering along any refusal-related direction produces nearly identical refusal--over-refusal trade-offs, acting as a shared one-dimensional control knob. Thus, different directions primarily affect not whether the model refuses, but how it refuses. Using sparse autoencoders, we uncover a structured internal representation of refusal: a reusable core of shared refusal latents supplemented by style- and domain-specific latents. Linear interventions collapse this structure into uniform behavioral control, flattening mechanistic differences across refusal types. Our results reconcile the apparent simplicity of refusal steering with the diversity of refusal behaviors, and clarify the limits of linear interpretability for aligned model behavior.

Read the original paper