Research
Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
Overview Research area: Mechanistic interpretability and representation geometry in transformer language models — specifically, why linear probes, sparse autoencoders (SAEs), and activation steering r

- arXiv
- 2602.09783
- Published
- 2026-02-10
- Authors
- Andres Saurez, Yousung Lee, Dongsoo Har
AI summary
Overview
Research area: Mechanistic interpretability and representation geometry in transformer language models — specifically, why linear probes, sparse autoencoders (SAEs), and activation steering recover stable semantic structure from nonlinear networks.
Technical level: Advanced. The paper is built around a formal theorem (with proofs in appendices), definitions of invariant subspaces and communicable features, and a capacity-constrained factorization argument, though the empirical sections and intuition (tokens "point" to their own features) are accessible without the formalism.
Scope: The paper argues that context-invariant linear feature directions are an architectural necessity of transformers, derives a "Self-Reference Property" from that result, and validates it across eight classification tasks and four model families (LLaMA3-8B, Mistral-7B, GPT2-Small, LLaMA3.2-3B).
What This Paper Is About
Linear interpretability methods such as linear probes and sparse autoencoders routinely extract meaningful semantic structure from transformer hidden states, but it is unclear why such simple tools should work inside deep nonlinear systems. The authors argue this is not just an empirical regularity but a structural consequence of architecture: transformers pass information through linear interfaces (the attention OV circuit and the unembedding matrix), so any feature read out through those interfaces must live in a context-invariant linear subspace. From this they derive that tokens themselves supply the geometric directions of their associated features, enabling zero-shot classification without labeled data or trained probes.
Key Contributions
-
An architectural explanation for linear interpretability. The authors prove the Invariant Subspace Necessity theorem: whenever a semantic feature is decoded through a linear interface, its representation must lie in a subspace shared across all contexts expressing that feature. This is positioned as complementary to work (Jiang et al., 2024) that explains linear representations via the next-token prediction objective and the implicit bias of gradient descent — optimization determines how representations are learned, architecture constrains what form they take.
-
The Self-Reference Property. Tokens directly provide the invariant direction for their own associated features (e.g., the token "France" serves as a reference vector for the France concept). This yields zero-shot identification of semantic directions and an unsupervised probe that classifies instances using only class token geometry, with no instance labels.
-
Convergent evidence for directional invariance. Sparse autoencoders trained only on implicit instance activations — with class tokens never introduced during training — still recover features that align with class token directions, indicating that both token-based probes and SAEs access the same underlying structure.
-
A capacity argument for feature sharing. Proposition 3.8 argues that because the vocabulary is much larger than the hidden dimension, token embeddings cannot occupy orthogonal directions and must factorize into shared feature directions, which then satisfy the conditions of the necessity theorem.
Main Findings
-
Invariant subspaces are required by linear interfaces (Theorem 3.7). For a feature decoded through a linear map, all contexts with the same feature value may differ only in directions orthogonal to the feature's readout vector, so the feature-relevant information must occupy a subspace determined by the feature and the interface, not by the individual context.
-
Linear versus nonlinear readouts change the geometry. Training a 2-layer transformer on modular division with p = 97 using an MLP classification head instead of a linear unembedding produced roughly 95% validation accuracy on the MLP head while a linear probe on the same hidden states failed at roughly 20%. Fourier structure (circular embeddings) appeared only in some random seeds, and precisely in those runs the linear probe succeeded — evidence that Fourier-like structure is permitted but not required under a nonlinear readout.
-
Tokens align with their class instances. Comparing class token directions to mean instance directions per attention head, across 89.6%–98.6% of attention heads the class token aligned more strongly with its own class than with other classes. Reported per-dataset head percentages above the diagonal: Countries 91.5%, Animals 97.6%, Cartoon Characters 86.0%, Emotions 89.6%.
-
Capacity forces factorization (Proposition 3.8). Because |V| >> d, token directions factorize as a sum of shared feature directions with |F| << |V|; the paper links this compression pressure to the information bottleneck principle.
-
The direction is invariant while the magnitude varies. Corollary 3.10 states that for any context expressing a directionally invariant feature, the hidden state decomposes as a context-dependent scalar times the invariant direction plus an orthogonal residual, which linear interfaces preserve.
-
All three probe families classify well (Table 2). On LLaMA3-8B, for Animals, SAE scored 92.05%, Unsupervised 94.32%, Zero-Shot 84.22%, and the Text Output baseline 99.67%; on Countries, SAE 75.60%, Unsupervised 79.60%, Zero-Shot 82.97%, Text Output 79.48%. GPT2-Small was markedly weaker (e.g., Animals: SAE 26.65%, Unsupervised 26.29%, Zero-Shot 25.39%, Text Output 35.33%). Supervised probes were omitted due to their tendency to overfit.
-
Learned transformations improve separability. The unsupervised probe consistently outperforms zero-shot classification, which the authors attribute to raw token directions sharing features (e.g., all country tokens encoding "being a country") that the contrastive objective disentangles.
-
SAEs recover the same directions as tokens. When top-k SAE dimensions of a class token (k = 32) were compared to the top-k most frequent SAE dimensions across that class's instances in the Animals dataset, the intersections were 15/32 for mammals, 23/32 for fish, 25/32 for reptiles, and 22/32 for birds.
-
Polysemy and partial synonymy are two sides of one coin. Combining the Fruits and Companies datasets, using a single "Apple" token gave 69% accuracy versus 65.7% when using separate tokens "Fruit apple" and "Company Apple". Both meanings coexist in superposition, with context modulating magnitude rather than selecting a direction; conversely "Dog" and "Cat" share a component along a mammal direction.
-
Generalization across model families. Consistent behavior across LLaMA3-8B, Mistral-7B, GPT2-Small, and LLaMA3.2-3B suggests directional invariance is not an artifact of a particular architecture or scale.
Methodology in Plain English
The authors start with a set of assumptions that hold for standard decoder-only transformers: the residual stream updates additively, attention OV circuits and the unembedding matrix act as linear maps on that stream, parameters are shared across token positions, and logits come from a linear projection. They then define what counts as a "communicable" feature — one that appears in multiple different contexts and can be recovered through a linear interface — and show by construction that such features must occupy a context-invariant subspace. Because a linear readout acts as a bottleneck, any context producing the same feature value must agree on its projection onto the feature's readout vector; the rest of the representation is free.
For practical use, they define the Identity-Projection operator: take the dot product of a hidden state with the normalized direction of a feature. The direction itself is obtained for free from the token that names the feature (the Self-Reference Property). On the empirical side, they build eight prompt datasets — Animals (6 classes, 50 instances each), Countries (5 classes, 39 each), Emotional Sentences (6, 60), Literary Quotes (6, 50), Cartoon Phrases (6, 50), Languages (6, 50), Fruits (4, 50), Companies (4, 50) — where each instance is an implicit sentence that expresses a class without naming it (e.g., "the country of the Eiffel Tower"). They then compare three probes that differ only in how representations are transformed before projection: a zero-shot probe using mean-centered class token directions (no training data), an unsupervised probe learning a transformation using only class tokens with a contrastive loss, and an SAE probe trained only on implicit instance activations, evaluated per attention head. A Text Output baseline scores each class by summed token log-likelihoods. Separately, they run the modular division experiment to test whether linear structure is driven by the readout interface rather than by training dynamics alone.
Why This Matters
Impact on research. The paper reframes linear interpretability from a lucky empirical regularity into a predicted consequence of transformer architecture. It offers a single explanation covering linear probes, sparse autoencoders, and direction-based steering, and it supplies a concrete diagnostic: instances that are not easily classifiable via token directions may indicate a feature that has not collapsed into an invariant form. It also gives a bridge between the token-direction and SAE literatures, and a way to identify which SAE features correspond to semantically meaningful directions.
- Label-free classification and auditing. The zero-shot and unsupervised probes classify instances using only class tokens, which could support rapid review of model behavior on new taxonomies without annotated datasets.
- Model editing and steering. Because token directions give a reference vector for a concept, the same geometry could guide activation steering toward or away from a feature.
- Feature dictionary evaluation. Token directions can act as a reference for judging whether SAE latents correspond to meaningful, reusable concepts.
- Diagnosing representation failure. Cases where the probe geometry breaks down flag concepts that the model has not represented in an invariant, linearly accessible way.
Industry relevance. Teams building or auditing large language models can obtain concept directions without training probes or labeling data, which lowers the cost of monitoring and steering model behavior. The results also suggest that SAE-based tooling and simpler token-direction baselines should be compared, since the paper reports both accessing the same structure. The authors' own Impact Statement declines to single out specific societal consequences beyond advancing machine learning.
Future Directions
- Scalable unsupervised circuit discovery. The Self-Reference Property suggests that tokens could anchor automated discovery of internal circuits without human-labeled concepts.
- New evaluation criteria for dictionary learning. The observed correspondence between SAE features and token directions could be turned into a test for whether a learned feature dictionary captures semantically meaningful structure.
- Characterizing when subspaces collapse. The authors explicitly do not characterize when invariant subspaces collapse to a single direction versus remaining higher-dimensional structures, leaving the conditions for directional invariance open.
- Features outside linear interfaces. Features operating through nonlinear gating, such as QK routing, may not exhibit the same invariance; the limits of the framework in these settings are unexplored.
Target Audience
Interpretability researchers and theorists studying why probing and steering methods work; machine learning practitioners applying SAEs, linear probes, or activation steering to language models; and graduate-level readers comfortable with linear algebra, subspace arguments, and transformer internals. Readers seeking applied recipes will find the zero-shot and unsupervised probe definitions immediately usable, while those interested in the theory will need the appendices for the full proofs.
Authors’ abstract
Linear probes and sparse autoencoders consistently recover meaningful structure from transformer representations -- yet why should such simple methods succeed in deep, nonlinear systems? We show this is not merely an empirical regularity but a consequence of architectural necessity: transformers communicate information through linear interfaces (attention OV circuits, unembedding matrices), and any semantic feature decoded through such an interface must occupy a context-invariant linear subspace. We formalize this as the \emph{Invariant Subspace Necessity} theorem and derive the \emph{Self-Reference Property}: tokens directly provide the geometric direction for their associated features, enabling zero-shot identification of semantic structure without labeled data or learned probes. Empirical validation in eight classification tasks and four model families confirms the alignment between class tokens and semantically related instances. Our framework provides \textbf{a principled architectural explanation} for why linear interpretability methods work, unifying linear probes and sparse autoencoders.