Research
Language Model Circuits Are Sparse in the Neuron Basis
Overview Research area: Mechanistic interpretability for large language models (natural language processing / machine learning). Technical level: Advanced. The paper assumes familiarity with Transform

- arXiv
- 2601.22594
- Published
- 2026-01-30
- Authors
- Aryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah Schwettmann
AI summary
Overview
Research area: Mechanistic interpretability for large language models (natural language processing / machine learning).
Technical level: Advanced. The paper assumes familiarity with Transformer internals, attribution methods, and sparse dictionary learning.
Scope: The paper argues and empirically demonstrates that MLP neurons themselves form a sparse, faithful basis for circuit tracing, removing the need for trained sparse autoencoders when locating causally effective subcomponents of a language model.
What This Paper Is About
Circuit tracing aims to localize a model's behavior to specific internal components, and most recent work assumes that individual neurons are polysemantic and uninterpretable, so it decomposes them into learned features such as sparse autoencoders (SAEs). The authors challenge that assumption, showing that the pre-down projection MLP activations are already as sparse a feature basis as SAEs. They then build an attribution pipeline that finds causally effective neurons directly, without any additional training cost.
Key Contributions
-
Empirical demonstration that MLP neurons are as sparse as SAEs. The authors state that they show for the first time that MLP neurons are as sparse a feature basis as SAEs, contradicting the common belief that neuron circuits are not sparse.
-
Two specific methodological changes that close the gap. They replace MLP outputs (post-down projection) with MLP activations (pre-down projection), a privileged basis, and they replace Integrated Gradients with RelP (Jafari et al., 2025), a single-backward-pass attribution method. The authors state they are the first to apply RelP directly to individual MLP neurons and to compute neuron-to-neuron edge weights.
-
An end-to-end gradient-based attribution pipeline for circuit tracing in the neuron basis. The pipeline surfaces causally effective neurons on multiple tasks and extends to edge-based circuit evaluation.
-
A replication-style case study on multi-hop reasoning. The authors obtain results in the neuron basis of Llama 3.1-8B-Instruct analogous to those reported by Lindsey et al. (2025) with cross-layer transcoders (CLTs), including neurons that can be steered to change the model's output.
Main Findings
-
MLP activations are far sparser than MLP outputs. On the subject-verb agreement (SVA) benchmark from Marks et al. (2025), MLP activations yield significantly smaller circuits than MLP outputs, by a factor of 100 times, and also close much of the gap with SAEs.
-
RelP closes the remaining gap with SAEs. With RelP, MLP activations reach near-perfect faithfulness and completeness with only approximately 200 neurons. RelP outperforms Integrated Gradients in almost all settings. Integrated Gradients is computed with 10 backward passes in their setup, while RelP requires only one backward pass.
-
The findings hold on unpaired data. When only original (not counterfactual) inputs are used for training, with a zero baseline for attribution and mean ablation for evaluation, the MLP activation neuron basis again requires a considerably smaller circuit than other methods. Results with zero ablation are reported in Appendix D.
-
RelP produces more faithful edges. For edge-based circuit evaluation on MLP activations, starting from the top 10^3 neurons (up to 5·10^5 potential edges per example), RelP with stop-gradients reaches over 80% faithfulness while maintaining high completeness with only approximately 10^5 edges, which is 10% of the candidate edges. The authors state it Pareto-dominates both alternatives (IG-inp. with 10 steps, and RelP without stop-grads on intermediate MLPs).
-
A small SVA circuit controls behavior. On the standard subject-verb agreement benchmark, a circuit of approximately 10^2 MLP neurons is enough to control model behavior.
-
The "Texas" circuit reproduces multi-hop reasoning structure. For the input "What is the capital of the state containing Dallas?", the pipeline recovered 257 total neurons with high attribution scores, from which the authors manually identified 23 neurons with particularly meaningful descriptions. These cluster into six groups matching categories from Lindsey et al. (2025): capital, state, Dallas, Texas, "say a capital", and "say Austin".
-
Cluster steering changes the model's prediction. All clusters except "state" change the model's top prediction when steered (with a single α in [-4, 4]). Suppressing the "Texas" cluster causes the model to still output state capitals, but for other states.
-
A single neuron can flip the answer. Steering L23/N8079-, described as "the phrase 'is' when referring to state capitals", with α in {0, 0.25, ..., 2} can flip the top output from the capital to the state in a majority of examples.
-
SAEs carry drawbacks the neuron basis avoids. The authors cite uninterpretable error terms from imperfect approximation, feature splitting and absorption leading to polysemanticity, and difficulty applying learned bases throughout training without additional computational overhead.
Methodology in Plain English
The authors frame circuit tracing as finding a small subgraph of the model's computation: a set of nodes (individual units such as MLP neurons or SAE features) and weighted edges capturing causal influence between them. They test several candidate bases: MLP activations, MLP outputs, attention outputs, and the residual stream, each in either the neuron basis or an SAE basis.
For the SVA experiments they use the Llama 3.1 8B base model and 8x width SAEs from Llama Scope. Each SVA task provides 300 training pairs of original and counterfactual inputs, with evaluation on 40 held-out pairs. The metric is the logit difference between the correct and incorrect verb form. Circuits are built by greedily taking the k highest-attribution nodes, and evaluated by mean-ablating the complement of the circuit and computing faithfulness and completeness (perfect is faithfulness 1 and completeness 0).
Attribution scores come from either Integrated Gradients (interpolating between counterfactual and original inputs) or RelP, which applies gradient-based attribution to a locally linear replacement model where nonlinearities such as RMSNorm, SiLU, and attention are replaced with frozen linear approximations. A "half rule" from layerwise relevance propagation divides the gradient by two through multiplicative interactions, preserving layer-by-layer conservation of attribution. Evaluations are always performed on the original model.
For the multi-hop case study they use Llama 3.1 8B Instruct on 50 state-capital questions, attribute from the sum of the top 5 next-token logits, filter nodes by an attribution threshold of 0.005 relative to the total logit value, and use RelP with stop-gradients for edges. Neurons are interpreted using automatic descriptions from Choi et al. (2024), where the best of 20 LM-generated descriptions was selected via a finetuned simulator. Steering multiplies a chosen set of neurons' activations by a scalar α and measures the change in output.
Why This Matters
Impact on research. The paper argues that the neuron basis has been under-explored in the circuits literature, so the uplift from SAEs is unknown. It recommends that new sparse dictionary-learning methods include an MLP neuron-level comparison. The authors note that Ameisen et al. (2025) do not report the minimal ablation of their gradient-based tracing algorithm with CLTs replaced by neurons, only a variant where neurons are thresholded by activation. Since the method requires no additional training, interpretability results can be obtained without training costs or loading SAEs into memory.
Real-world applications (as implied by the paper's framing):
- Auditing and overseeing AI systems by inspecting internal reasoning steps that the model does not verbalize, both at inference time and during training.
- Controlling model behavior through steering, as demonstrated by flipping outputs between capitals and states in the case study.
- Data and model diffing, and comparing behavior across checkpoints without the computational overhead of learned bases.
- Enabling independent auditors and democratizing AI safety research through a more tractable, scalable tracing technique.
Industry relevance. Because the approach avoids training auxiliaries and avoids holding SAEs in memory, it is described as tractable for large models. The authors also note architectural trends favoring MLP sparsity: the shift to gated MLPs may allow packing more features without sacrificing sparsity, and mixture-of-experts MLPs explicitly enforce sparsity by routing inputs to a subset of expert MLPs.
Caveats the authors raise. The paper states that this tool may also enable intervening on model internals to produce harmful outputs, though the authors judge the research societally beneficial on balance.
Future Directions
-
Clustering and describing neurons at scale. The authors state there are still too many neurons in a reasonably comprehensive circuit to interpret easily, and that principled approaches for clustering neurons plus better natural-language descriptions of neurons and clusters are needed.
-
Efficiency of the attribution implementation. The current implementation relies on serial calls to
torch.autograd.grad, leading to low compute utilization. The authors point to Arora et al. (2026) for an end-to-end automated circuit tracing pipeline that addresses this to a great extent. -
Understanding the true feature basis. The authors argue that until the properties of the true feature basis are better understood, it is prudent to exhaust the neuron basis first, and that sparse dictionary methods make assumptions about feature geometry and frequency with unclear benefits. They note alternative techniques without reconstruction error are compatible with their approach.
-
Extending the case studies. The paper reports three additional tasks in Appendix H, and mentions new results on user modelling, suggesting room to widen the set of behaviors traced in the neuron basis.
Target Audience
Mechanistic interpretability researchers and AI safety practitioners who build or evaluate circuit-tracing methods and who currently default to SAEs, transcoders, or CLTs. It is also relevant to engineers who need an interpretability pipeline that runs without training auxiliary models or loading large dictionaries into memory, and to researchers comparing attribution techniques such as Integrated Gradients against RelP.
Authors’ abstract
The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986). Language model interpretability research has thus turned to techniques which decompose the neuron basis into more interpretable units of model computation, such as sparse autoencoders (SAEs). However, not all neuron-based representations are uninterpretable. For the first time, we empirically show that MLP neurons are as sparse a feature basis as SAEs. We use this finding to develop an end-to-end gradient-based attribution pipeline for circuit tracing on the MLP neuron basis, which surfaces causally effective neurons on a variety of tasks. On a standard subject-verb agreement benchmark (Marks et al., 2025), a circuit of $\approx 10^2$ MLP neurons is enough to control model behaviour. On the multi-hop city-state-capital task from (Lindsey et al., 2025), we find a circuit in which small sets of neurons encode specific latent reasoning steps (e.g. mapping a city to its state), and can be steered to change the model's output. This work thus advances automated interpretability of language models without imposing additional training costs.