Research
Rank-1 LoRAs Encode Interpretable Reasoning Signals
Rank-1 LoRAs Encode Interpretable Reasoning Signals Overview Research area: Mechanistic interpretability of large language models, specifically parameter-efficient fine-tuning (LoRA) and sparse autoen

- arXiv
- 2511.06739
- Published
- 2025-11-10
- Authors
- Jake Ward, Paul Riechers, Adam Shai
AI summary
Rank-1 LoRAs Encode Interpretable Reasoning SignalsOverview
Research area: Mechanistic interpretability of large language models, specifically parameter-efficient fine-tuning (LoRA) and sparse autoencoders applied to reasoning models.
Technical level: Advanced. The paper assumes familiarity with LoRA adapters, low-rank matrix decomposition, sparse autoencoders, activation steering, KL divergence, and chain-of-thought reasoning benchmarks.
Scope: The paper shows that a rank-1 LoRA adapting every projection matrix of Qwen-2.5-32B-Instruct recovers 73-90% of the reasoning-performance gain of a full-parameter finetune, and that the resulting adapter activations can be interpreted with LLM autointerpretation and a cross-layer sparse autoencoder.
What This Paper Is About
Reasoning models gain large performance improvements on logical tasks through inference-time chain-of-thought, but the parameter changes responsible for those gains are numerous and diffuse, making white-box interpretation difficult. This paper sidesteps that problem by deliberately constraining the finetune to a single rank, so the difference between base and finetuned model becomes a small, inspectable set of scalar activations. The goal is both to show that minimal adapters can elicit reasoning performance and to use those adapters as a targeted lens for understanding what reasoning behavior looks like inside the network.
Key Contributions
- The authors train and open-source a rank-1 LoRA that adapts all projection matrices in Qwen-2.5-32B-Instruct, and show it recovers 73-90% of reasoning-benchmark performance compared to a full-parameter finetune while having less than 0.03% as many trainable parameters as the base model.
- They show that individual adapter activations are monosemantic roughly as often as MLP neurons, and interpret those directions, finding they fire more often on reasoning-specific categories such as answers or solutions, problem instructions, mathematical symbols, and reasoning discourse.
- They train a cross-layer sparse autoencoder on the entire 448-dimensional LoRA activation state and identify fine-grained, monosemantic features that organize into categories such as Mathematical Operators, Procedural Markers, and Discourse and Reasoning Markers.
- They perform component-wise and simultaneous ablation studies, finding that MLP adapters (especially those on
gate_proj) drive most of the behavioral change, with mid-to-late layers mattering most.
Main Findings
- Benchmark recovery: The rank-1 LoRA reaches 0.5000 on AIME'24 (no-figures) versus 0.2333 for the base model and 0.6000 for the full finetune (72.73% recovery); 0.9100 on MATH500 versus 0.8340 base and 0.9220 full finetune (86.36% recovery); and 0.5808 on GPQA-Diamond versus 0.4899 base and 0.5909 full finetune (89.90% recovery).
- Adapter scale: The LoRA encodes 192 MLP and 256 attention adapter components across layers and weight matrices, giving a total 448-dimensional activation state, one scalar per adapter component per token position.
- Comparable monosemanticity to neurons: LoRA activations have roughly the same likelihood as MLP neurons to be monosemantic, measured against a baseline of the first 60 neurons of each MLP in the unadapted base model.
- Different feature profile: Relative to MLP neurons, LoRA activations are more likely to fire for text corresponding to answers or solutions, problem instructions, mathematical symbols, and reasoning discourse.
- SAE on the whole adapter state: A batch-top-k SAE with k=16 and an expansion factor of 8, trained on the 448-dimensional LoRA activation state, yields roughly 2000 features after filtering dead latents.
- Improved monosemanticity from the SAE: LLM-based classification identifies 62% of SAE features as "cleanly monosemantic," up from 22% of LoRA features, with an additional 22% of SAE features classified as "broad but consistent."
- Feature concentration: SAE features tend to concentrate on mathematical operators and syntax, numbers, formatting tokens, and reasoning control-flow. The authors note the categorization methodology is subjective and sensitive to prompt verbiage, and present the visualization only as a tool for high-level intuitions.
- Layer and component importance: Individual component ablation shows mid-to-late layers (particularly 44, 45, 46, and 62) have the greatest effect on downstream KL divergence relative to the unmodified LoRA, and MLP components have significantly larger impact than attention components, with
gate_projhaving the strongest average effect. - Simultaneous ablation: Removing all attention adapters decreases performance but still outperforms the base model, while removing all MLP adapters causes severe degradation, underperforming the base model on one of three tasks. Concretely, on AIME'24 the full LoRA scores 0.5000 (100%), attention-ablated 0.3667 (50.02%), and MLP-ablated 0.1333 (-37.50%); on MATH500, 0.9100 (100%), 0.9000 (86.84%), and 0.8440 (13.16%); on GPQA-Diamond, 0.5808 (100%), 0.5152 (27.83%), and 0.5051 (16.72%).
- Cherry-picked interpretable directions: The appendix shows adapter directions that consistently activate on the token "Wait" (and surrounding commas/periods), on single-letter math variables and exponent tokens, and on positional numbering in molecules.
- Cherry-picked SAE features: Examples include hesitation markers such as "Hmm" in reflective internal thought, firing on "is" as an equality indicator in mathematical definitions, sphere volume formula tokens (specifically the 4/3, pi, and R^3 components), and firing on "closed" in mathematical topology contexts.
Methodology in Plain English
The researchers took a base instruction-tuned model and finetuned it on a small dataset of reasoning traces, but with a key restriction: instead of updating all the model's weights, they attached a rank-1 adapter to every MLP matrix (up_proj, down_proj, gate_proj) and every attention matrix (Q, K, V, O) in every layer. Because the rank is 1, each adapter is just two vectors, and the value flowing through it during a forward pass is a single number per token. That single number is the unit of analysis.
Training used s1k-1.1, a sample-efficient dataset of 1000 DeepSeek R1 chain-of-thought trajectories and answer attempts across diverse reasoning problems, for 5 epochs with cross-entropy loss on a single 8xH200 node. The authors then ran the model over the training set, recorded the 448 adapter activations for every token, and interpreted them in two ways: directly, by extracting max-activating examples for each adapter direction and comparing them against a baseline of MLP neurons; and collectively, by training a sparse autoencoder on the full 448-dimensional activation vector to find sparser, more monosemantic features. Interpretation and categorization were automated with LLMs — gpt-5-mini for generating interpretations and classifying monosemanticity, and claude-opus-4.1 for generating the high-level feature categories from a mixed list of MLP neuron and LoRA direction interpretations.
To find out which parts of the adapter actually matter, they ablated components: zeroing out individual components and whole layers to measure KL divergence from the unmodified LoRA's output distribution, and separately removing all MLP adapters or all attention adapters at once to measure benchmark impact.
Why This Matters
The work reframes how interpretability researchers can study emergent capabilities. Instead of trying to diff two fully finetuned models where parameter changes are "many and diffuse," it enforces minimality in parameter space so the change becomes tractable, then reads off interpretable signals. It also suggests that reasoning performance can arise largely from minimal changes to base model parameters rather than requiring widespread rewiring, and that parameter-efficient training can double as an experimental instrument rather than just a compute-saving trick.
Real-world applications:
- Cheaper domain adaptation: If a rank-1 adapter recovers most of the gain of a full finetune, practitioners can specialize models for reasoning-heavy domains with dramatically smaller checkpoints and training costs.
- Safety and auditing: Interpretable adapter directions and SAE features give auditors a small, inspectable surface for checking what a finetune actually changed, rather than auditing a full weight delta.
- Adapter composition and debugging: Since MLP adapters — especially
gate_proj— dominate the effect, engineers building multi-adapter systems know where to look when adapters conflict or fail. - Education and tooling: Named, monosemantic features such as "Wait" hesitation markers or equality indicators can power dashboards that visualize model reasoning for human inspection.
Industry relevance: parameter-efficient finetuning is already standard in production; this paper argues the artifacts it produces are also interpretability assets, which matters for teams that need to ship specialized reasoning models while meeting audit and safety requirements.
Future Directions
- Identifying the circuits LoRAs interface with. The authors explicitly propose focused analysis of which existing base-model circuits the adapter directions connect to, hoping this illuminates the core computational mechanisms behind reasoning.
- Resolving the steering question. The authors tried to show a causal effect by using extracted LoRA directions for steering but got inconclusive results: effects required magnitudes in excess of 50x normal activation magnitudes, at which point the model showed a high propensity for backtracking and emitting tokens like "Wait." More investigation is needed before causal claims can be made.
- Scaling beyond one model and one benchmark family. The study only covers Qwen-2.5-32B-Instruct and evaluates mostly on math-related reasoning benchmarks, so replication on other models and non-math reasoning tasks is open.
- Testing whether the effect holds with larger training data. The LoRA was trained on roughly 10 million tokens; the authors note the model is likely missing circuits that training on larger datasets would produce, raising the question of how much of the reasoning capability stays recoverable at rank 1 as base-model training grows.
Target Audience
Mechanistic interpretability researchers, especially those working on sparse autoencoders, activation steering, and feature decomposition. Also useful for engineers doing parameter-efficient finetuning who want to know how much of a full finetune's behavior a rank-1 adapter can preserve and which components carry the load, and for safety researchers looking for small, auditable surfaces on which to inspect model behavior changes.
Authors’ abstract
Reasoning models leverage inference-time compute to significantly enhance the performance of language models on difficult logical tasks, and have become a dominating paradigm in frontier LLMs. Despite their wide adoption, the mechanisms underpinning the enhanced performance of these reasoning models are not well understood. In this work, we show that the majority of new capabilities in reasoning models can be elicited by small, single-rank changes to base model parameters, with many of these changes being interpretable. Specifically, we use a rank-1 LoRA to create a minimal parameter adapter for Qwen-2.5-32B-Instruct which recovers 73-90% of reasoning-benchmark performance compared to a full parameter finetune. We find that the activations of this LoRA are as interpretable as MLP neurons, and fire for reasoning-specific behaviors. Finally, we train a sparse autoencoder on the entire activation state of this LoRA and identify fine-grained and monosemantic features. Our findings highlight that reasoning performance can arise largely from minimal changes to base model parameters, and explore what these changes affect. More broadly, our work shows that parameter-efficient training methods can be used as a targeted lens for uncovering fundamental insights about language model behavior and dynamics.