Skip to content
AI.info

Research

Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential

Overview Research area: Mechanistic interpretability of large language models, specifically the internal circuitry of refusal behavior and how jailbreaks bypass it. Technical level: Advanced. The pape

Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential
arXiv
2610.05541
Published
2026-10-04
Authors
Swadesh Swain, Sanghamitra Dutta

AI summary

Overview

Research area: Mechanistic interpretability of large language models, specifically the internal circuitry of refusal behavior and how jailbreaks bypass it.

Technical level: Advanced. The paper assumes familiarity with sparse autoencoders, transcoders, attribution graphs, ablation, and statistical effect sizes.

One-sentence scope: The paper introduces a metric (Counterfactual Activation Potential) and a discovery algorithm (CSFD) for finding safety-critical model features that are silent on a prompt because other features actively suppress them, and shows these suppressed features are causally responsible for refusals.

What This Paper Is About

Existing interpretability tools for LLM safety look only at which neurons or features are active on a given prompt. Roughly 99% of transcoder features are silent on any prompt, and the paper argues this silent majority is not uniform: some features are silent only because the input does not drive them, while others are driven but are held below threshold by negative influence from neighboring features. The goal is to define, score, and efficiently discover these "suppressed" safety features, and to show that jailbreaks can succeed by suppressing them rather than by activating anything harmful.

Key Contributions

  1. Evidence that jailbreaks suppress safety-critical features. The authors show that zeroing a single validated feature flips refusals into compliance at rates up to 82.5%, that the PAIR jailbreak attack naturally raises suppression on the highest-scoring features, and that amplifying a feature's suppressors (with the prompt unchanged) lowers its activation and raises compliance.

  2. The Counterfactual Activation Potential (CAP) score. CAP is defined as the product of three components for a validated feature: encoder alignment or drive (z_alpha), suppression strength (z_gamma), and safety criticality (SC). Because it is a product, it is large only when all three hold simultaneously, and unlike a sum it cannot be rescued by one large term compensating for smaller ones. The two continuous terms are standardized within a model; SC is left unnormalized because it is already a fraction.

  3. The CAP-guided Safety Feature Discovery (CSFD) algorithm. A two-stage pipeline that avoids exhaustive ablation over millions of features. Stage 1 filters using Expected Activation Differential (the difference in "activation mass" — mean activation times activation count — between refused and benign prompts), then re-ranks survivors by Cohen's d, using statistics already recorded during inference. Stage 2 validates the reduced candidate set by ablation.

  4. Experimental validation across five models. Gemma-2-2B, Qwen3-8B, Gemma-3-4B-IT, Qwen3-1.7B, and Llama-3.2-1B, using approximately 86K prompts from WildJailbreak, with an independent replication on HarmBench.

Main Findings

  • Candidates are overwhelmingly causally valid. Of 50 CSFD candidates per model on WildJailbreak, validity (SC > 0) was 47 of 50 (94%) for Gemma-2-2B, 42 of 50 (84%) for Qwen3-8B, 39 of 50 (78%) for Gemma-3-4B-IT, and 50 of 50 (100%) for both Qwen3-1.7B and Llama-3.2-1B. Maximum SC per model was 0.81, 0.538, 0.825, 0.791, and 0.733 respectively.

  • Zeroing one feature flips up to 82.5% of refusals. Reported ablation flip rates reach 81.0% on Gemma-2-2B, 82.5% on Gemma-3-4B-IT, and 79.1% on Qwen3-1.7B.

  • Natural jailbreaks suppress these features. Under PAIR rewriting, suppression on the highest-CAP features rises (2–4× in the abstract's summary) and their activation falls by up to 80%. On Qwen3-8B, L31/F97078 falls from 44.1 to 19.9 at g_f = 3.3 and L35/F133412 from 51.7 to 20.9 at g_f = 3.2. The largest effect is Qwen3-1.7B's L27/F37712, whose suppression quadruples and whose activation falls from 59.6 to 11.9. Across the three instruction-tuned models, activation falls only when suppression rises (g_f > 1), with one exception, L33/F651 on Gemma-3-4B-IT.

  • Suppression alone is sufficient to cause compliance. Scaling a feature's suppressors by 2, 4, and 8 times lowers the feature's own activation by 24%, 35%, and 42% on average and raises the flip rate from zero to 0.09, 0.18, and 0.29. The same scaling of random active features from the same layers moves activation by 6% and flips 0.05 of prompts.

  • Reversal is asymmetric. Scaling suppressors down on a jailbroken prompt, or restoring the feature to its activation on the original refused prompt, rarely reinstates refusal — indicating refusal is distributed across several features rather than gated by one.

  • CAP separates harmful from benign prompts better than baselines. AUROC for benign vs. harmful-refused: 0.788 (CAP probe) vs. 0.546 (DIM) vs. 0.668 (raw activations) on Gemma-2-2B; 0.719 vs. 0.453 vs. 0.421 on Qwen3-8B. Raw activations of the same 50 candidates reach only 0.657 on Gemma-2-2B.

  • DecAlign misses these features. On Gemma-2-2B, candidates point toward refusal (mean cosine +0.033 vs. −0.003 for random features). On Qwen3-8B they point away from refusal: CAP is anti-correlated with the DecAlign score at Spearman ρ = −0.818, and ranking the ten Qwen3-8B features by DecAlign recovers only one suppressed feature among its top five, versus four among the five highest-CAP features.

  • All three CAP terms are needed for ranking, but not for probing. Full CAP predicts alignment with the jailbreak direction at ρ = +0.661 (p = 0.037); no single term reaches significance (SC +0.539, p = 0.108; z_alpha +0.430, p = 0.215; z_gamma +0.442, p = 0.201). Removing SC leaves the probe AUROC unchanged on Gemma-2-2B, Qwen3-1.7B, and Llama-3.2-1B.

  • Specialization tracks criticality. Among ten interpreted Gemma-2-2B features, L25/F13235 (classified, medical, or confidential information) and L24/F4006 (short imperative instructions for harmful actions) have the highest SC at 0.852 and 0.824, while the broadest feature, L25/F14180, has the lowest at 0.490.

  • Filtering is fast and stable. On Qwen3-8B the filter starts from 35,831 features firing on at least one prompt; the coverage floor retains 3,015, positive EAD retains 1,670, mass ratio retains 1,080, and the top 50 by Cohen's d form the candidate set. On Gemma-2-2B, 107 of 2,301 active features survive. Both runs complete in under a minute. Candidates concentrate in late layers (Qwen3-8B: layers 29–35, with 52% in layer 35 alone), except Gemma-2-2B, whose validated features span layers 9 to 25, with L9/F65 reaching SC = 0.65.

  • Alignment correlates with distribution. Qwen3-8B refuses 36.8% of WildJailbreak prompts versus Gemma-2-2B's 22.9%, and is jailbroken on 23.4% versus 27.1%; its peak SC is correspondingly lower (0.538 vs. 0.81), with mean SC 0.356 vs. 0.716 over each model's ten features. Judge agreement on flip labels ranges from 93.6% to 96.5% across three LLM judges.

Methodology in Plain English

The authors work with per-layer transcoders (PLTs): dictionaries of thousands or millions of interpretable features per layer, each with an encoder vector that decides whether the feature fires and a decoder vector that writes its output into the model's residual stream. A feature is active if its activation exceeds a threshold at some token position.

They formalize attribution edges between features — a positive edge means one feature promotes another, a negative edge means it suppresses it. The suppressors of feature f on prompt P are the active features with a negative edge into f above a magnitude threshold τ. A feature is called suppressed if its encoder alignment is positive but its suppressors are active, so suppression rather than lack of drive explains its silence.

To find such features without ablating everything, Stage 1 of CSFD uses activation statistics already computed during inference: it keeps features with positive Expected Activation Differential, a minimum number of refused-prompt activations, and a mass ratio above a selectivity floor, then ranks survivors by Cohen's d. Stage 2 then ablates each candidate (clamping its activation to zero from the last input token onward) on harmful-refused prompts where it is active, using greedy decoding with at most 512 new tokens and a Gemini 3.0 Flash judge at temperature 0.1 to label responses as benign, harmful refused, or jailbreak success.

Validation proceeds through three experiments: (i) ablation to measure safety criticality, the fraction of refused prompts that flip to compliance; (ii) matched-pair comparison of the same request before and after PAIR turns it into a successful jailbreak, measuring the change in suppression ratio g_f and in activation; and (iii) direct intervention scaling a feature's suppressors by a factor k while leaving the prompt untouched.

Why This Matters

Impact on research. The paper reframes jailbreaks as an act of suppression rather than activation, and argues that activation-focused interpretability is structurally blind to a relevant

Authors’ abstract

Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it). To find suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. A significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. Under natural jailbreaks, the suppression acting on the highest-CAP features rises 2-4x, and their activation correspondingly falls by up to 80%. Amplifying a feature's suppressors pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experiments span five Gemma, Qwen, and Llama models across various parameter sizes. Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.

Read the original paper