Skip to content
AI.info

Research

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

Overview Research area: Mechanistic interpretability and self-interpretation for large language models, sitting at the intersection of activation patching, sparse autoencoders, and parameter-efficient

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
arXiv
2602.10352
Published
2026-02-10
Authors
Keenan Pepper, Alex McKenzie, Florin Pop, Stijn Servaes, Martin Leitgab, Mike Vaiana, Judd Rosenblatt, Michael S. A. Graziano, Diogo de Lucena

AI summary

Overview

Research area: Mechanistic interpretability and self-interpretation for large language models, sitting at the intersection of activation patching, sparse autoencoders, and parameter-efficient adaptation.

Technical level: Intermediate. The core idea is intuitive, but following the architecture comparisons and evaluation protocols requires familiarity with residual stream activations, Patchscopes, and SAE feature labeling.

Scope: The paper shows that a frozen language model can be made to describe its own internal activations reliably by training a very small adapter (as few as d_model + 1 parameters) on existing vector-label interpretability artifacts, rather than fine-tuning the model or hand-tuning injection scales.

What This Paper Is About

Methods that ask a language model to describe its own internal states ("self-interpretation") are fragile: the injected activation usually needs a carefully tuned scale, and most scales produce fluent but ungrounded explanations. The authors' goal is to replace manual scale tuning with a learned transformation, trained on data that interpretability research has already produced, such as sparse autoencoder (SAE) decoder vectors paired with labels and contrastive activation vectors paired with topic descriptions. Crucially, they keep the language model itself entirely frozen.

Key Contributions

  1. Training lightweight adapters instead of fine-tuning the model. The authors learn a mapping function f(h) between an activation vector and the model's token embedding space, using existing (vector, label) interpretability artifacts as supervision. The language model weights never change, so the interpreter and the subject model remain identical.

  2. Demonstrating that an extremely small adapter suffices. A scalar affine adapter with only d_model + 1 parameters captures most of the gain. The learned bias vector alone accounts for roughly 85% of the loss improvement over untrained baselines, and adding a low-rank term yields further but diminishing gains.

  3. Characterizing when adapter capacity helps and when it hurts. Full-rank affine adapters succeed on Wikipedia contrastive topic vectors but overfit catastrophically on SAE features, a difference the authors trace to the intrinsic dimensionality of the two representation classes.

  4. Showing self-interpretation scales with model size faster than general capability. Using a "Taboo" baseline as a capability ceiling on Qwen-2.5 models (7B, 14B, 32B, 72B), the gap between trained self-interpretation and the ceiling narrows with scale.

Main Findings

  • Bias vector dominates. Scale-only adapters improve validation loss by 0.291 over identity, while adding the bias vector (scalar affine) yields an additional 2.75 improvement, and the single d-dimensional bias accounts for approximately 85% of the total gain from the best adapter.

  • Scalar affine is a strong minimal baseline. With 4097 parameters (d = 4096 for Llama-3.1-8B), scalar affine reaches validation loss 1.787 on Llama Scope SAEs, with essentially zero train-val gap.

  • Low-rank additions help, with diminishing returns. SA + LR at rank 64 achieves the best validation loss of 1.619; rank 256 gives 1.622.

  • Full-rank overfits on SAE data but wins on contrastive vectors. Full-rank affine (16.8M parameters) reaches validation loss 1.743 with a train loss of 0.64, while on Wikipedia contrastive vectors it achieves 82.9% recall@1 versus 0.04% for untrained SelfIE without scale tuning, and 98.4% versus 0.9% at recall@100.

  • Contrastive vectors have low intrinsic dimensionality. Wikipedia contrastive vectors concentrate over 90% of their variance in roughly 200 dimensions, which implicitly regularizes the full-rank transformation; SAE features span nearly the full activation space.

  • Adapters can beat their own training labels. At 70B scale, trained adapters reach 70% generation scoring versus 50% for the training labels. On Llama-3.3-70B-Instruct, the SA+LR adapter achieves a 69.7% mean hit rate on held-out Goodfire latents, compared to 50.0% for repeated auto-interpretation labels, 60.4% for paraphrased labels, and 48.1% for untrained SelfIE.

  • Cross-dataset generalization favors simpler adapters. The Wikipedia-trained scalar affine adapter reaches 39.3% and 47.0% generation scoring hit rates on SAE latents despite never training on SAE features, while the Wikipedia-trained full-rank adapter reaches only 22.1% and 32.0%.

  • Detection and generation scoring disagree. The SA+LR adapter trained on Llama Scope achieves the highest detection F1 (0.722), tied with the auto-interp baseline (0.722) and above paraphrases (0.712), while scalar affine wins on generation scoring. At 70B, SA wins on detection (F1 = 0.760) while SA+LR wins on generation.

  • Comparable performance to LoRA without modifying weights. Rank-4 LoRA Activation Oracles achieve 81.6% ±0.6 recall@1 on Wikipedia topics versus 82.9% ±0.6 for trained SelfIE, and 62.3% ±0.6 hit rate on Goodfire SAE generation scoring versus 59.2% ±0.6.

  • Qualitative examples show semantic, not surface-level, recovery. On a held-out prompt about "propagating gradients back through a neural network," five generations at temperature 0.5 all converged on backpropagation or automatic differentiation, without the prompt using the word "backpropagation."

  • Self-interpretation improves faster than the capability ceiling. Untrained SelfIE stays below 2% recall@100 at all Qwen-2.5 scales, while trained SelfIE closes the gap to the Taboo baseline as models grow from 7B to 72B.

  • Implicit reasoning can be decoded. On 500 sampled TwoHopFact prompts where Llama-3.1-8B-Instruct answered both hops correctly with no chain of thought, the trained adapter detected the bridge entity in 455 cases (91.0% ±1.3) versus 282 (56.4% ±2.2) for untrained, a 4.8x reduction in undetected bridge entities. Only 2 of 500 prompts showed the reverse pattern.

  • SelfIE outperforms linear probes while producing free-form language. At n = 10 generations per position, SelfIE achieves 73.0% bridge entity detection at its best layer versus 67.0% for the best probe layer; pooled across all layers and positions, SelfIE achieves 91.0% ±1.3 versus 85.6% for probes, using one adapter against 32 probes.

Methodology in Plain English

The approach reuses the Patchscopes framework. The researchers take an activation vector from inside a language model, transform it with a learned function, and inject the result at the placeholder position of an explanation-seeking prompt such as "What is the meaning of 'TOKEN'? The meaning of 'TOKEN' is." The model then generates a description autoregressively.

Rather than hand-tuning the injection strength, the authors learn the transformation from existing interpretability artifacts. They use two kinds of training pairs. The first is SAE decoder vectors paired with auto-interpretability labels: 45,418 Goodfire decoder vectors from a layer 19 residual stream SAE on Llama-3.1-8B-Instruct, 61,521 features from a layer 50 SAE on Llama-3.3-70B-Instruct, and Llama Scope SAE decoder vectors from 32k- and 131k-width SAEs trained on layers 0-31 of Llama-3.1-8B. The second is 49,637 Wikipedia "Vital Level 5" article titles turned into contrastive vectors at layer 19, each paired with approximately 15 to 20 synthetic labels.

Training is ordinary supervised learning: minimize cross-entropy on the label tokens after injecting f(h), with the language model frozen. All input vectors are normalized to unit L2 norm, and a separate scale factor controls injection magnitude at inference. At evaluation time they generate N = 6 candidate descriptions per vector by varying the injection scale over a logarithmic grid, using independent scoring seeds for selection and reporting to avoid selection-on-noise effects. Contrastive vector results are scored by embedding-based retrieval with GTE-large, and SAE results use detection scoring and generation scoring.

Why This Matters

Impact on research. Mechanistic interpretability has already produced large collections of labeled directions. This paper reframes those artifacts as training data rather than endpoints of analysis. Because the base model is frozen, the interpreter is identical to the subject, preserving the "privileged access" hypothesis maximally, in contrast to fine-tuning approaches such as LatentQA, Activation Oracles, and encoder-decoder sparse bottleneck methods. The authors also note that scalar affine adapters have far fewer degrees of freedom for learning spurious input-dependent patterns than fine-tuned decoders do.

Real-world applications:

  • Alignment auditing. Surfacing latent states a model acts on but never verbalizes, which the paper describes as "catching models red-handed," as when the bridge entity "Plato" is recovered from a prompt whose answer is "Athens."
  • Model transparency and detection of hidden objectives or behavioral changes. The authors point to concurrent work that recovers secrets fine-tuned into models and detects jailbreaks, and suggest lightweight adapters could target similar auditing tasks with extensions.
  • Feature labeling at scale. Generating more accurate and more consistent feature descriptions than the noisy auto-interpretability labels the adapters were trained on, which matters for anyone depending on SAE feature libraries.
  • Cost-efficient interpretability pipelines. Training reportedly requires roughly 10 GPU-hours at 70B scale and needs no SAE at inference time beyond the source activation vector.

Industry relevance. The method requires only vector-label pairs that frontier labs already produce in abundance (the paper cites millions of labeled SAE features), and the adapter is trained once and reused across inference tasks. Because the served model is untouched, the technique offers a low-risk way to add interpretability tooling to an existing deployment.

Future Directions

  • Combining adapters with fine-tuning approaches. The authors state that decoder fine-tuning and trained adapters are complementary and could be combined, but leave this to future work. Adopting the Q&A framing and diverse training objectives from concurrent work, particularly self-supervised context prediction, could expand capabilities while keeping the adapters simple.

  • Better training objectives. The scale minimizing cross-entropy often differs from the scale maximizing generation scoring, suggesting future work on objectives that directly optimize interpretation quality rather than label likelihood.

  • Verifiable self-interpretation and RL from internal rewards. The authors propose making a model's claims about its internals testable via generation scoring, then using those testable claims as training signal to optimize models for accurately reporting their own computations.

  • Extending to the auditing applications demonstrated by weight-modifying methods. Recovering hidden objectives and detecting behavioral changes from fine-tuning are described as compelling targets that lightweight adapters might address with appropriate extensions.

Target Audience

This paper is most useful to interpretability and AI safety researchers who work with activation patching, sparse autoencoders, or feature labeling; to alignment engineers interested in auditing model internals without altering deployed weights; and to practitioners who want an inexpensive way to turn existing labeled feature collections into a working self-interpretation pipeline. Readers without background in transformer internals will need supporting material on residual stream activations and the Patchscopes framework to follow the architecture comparisons and evaluation sections.

Authors’ abstract

Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitivity. We show that training lightweight adapters on interpretability artifacts, while keeping the LM entirely frozen, yields reliable self-interpretation across tasks and model families. A scalar affine adapter with just $d_\text{model}+1$ parameters suffices: trained adapters generate sparse autoencoder feature labels that outperform the training labels themselves (70% vs 50% generation scoring at 70B scale), identify topics with 94% recall@1 versus 1% for untrained baselines, and decode bridge entities in multi-hop reasoning that appear in neither prompt nor response, surfacing implicit reasoning without chain-of-thought. The learned bias vector alone accounts for 85% of improvement, and simpler adapters generalize better than more expressive alternatives. Controlling for model knowledge via prompted descriptions, we find self-interpretation gains outpace capability gains from 7B to 72B parameters. Our results demonstrate that self-interpretation improves with scale, without modifying the model being interpreted.

Read the original paper