Research
Equivariant Sparse Autoencoders: Mechanistic Interpretability of Neural Networks on Symmetric Data
Overview Research area: Mechanistic interpretability (MI) of neural networks, specifically sparse autoencoders (SAEs), combined with equivariant/geometric deep learning. The authors are Ege Erdogan an

- arXiv
- 2511.09432
- Published
- 2025-11-12
- Authors
- Ege Erdogan, Ana Lucic
AI summary
Overview
Research area: Mechanistic interpretability (MI) of neural networks, specifically sparse autoencoders (SAEs), combined with equivariant/geometric deep learning. The authors are Ege Erdogan and Ana Lucic at the University of Amsterdam (arXiv:2511.09432v2, cs.LG).
Technical level: Advanced. The paper assumes familiarity with group theory, sparse dictionary learning, the Linear Representation Hypothesis, and probing-based evaluation of neural network representations.
Scope: The paper extends the Linear Representation Hypothesis to account for input symmetries, designs "Equivariant SAEs" that decompose activations into invariant features plus a learned linear transformation, and evaluates them on one toy model, one synthetic image dataset, and two real scientific image datasets across five base models.
What This Paper Is About
SAEs disentangle dense neural network activations into sparse, interpretable features, but they suffer from unidentifiability: different explanations can fit the data equally well without being more faithful to the underlying model. The authors show this problem is worsened by data symmetries such as rotations, which are common in scientific domains, because a feature for a given concept can appear as a different direction depending on orientation. Their goal is to build SAEs whose priors match the known symmetries of the data, and to test whether such symmetry-aware SAEs recover and expose more useful features than standard SAEs.
Key Contributions
- The authors extend the Linear Representation Hypothesis (LRH) to settings with group symmetries, reformulating it so that activations of a transformed input equal a linear transformation ρ(g) applied to the sparse sum of features of the canonical input (Equation 2).
- They design Equivariant SAEs, consisting of a group-invariant SAE plus separately learned linear transformation matrices that model how activations change under input transformations, with three candidate invariance losses (Canonical, Latent, Output) and a two-step training procedure.
- They show on a toy model and on neural network activations from three datasets and five base models that Equivariant SAEs improve feature recovery/detection and downstream probing usefulness, with the largest gains when the base model's activations vary more under input rotations.
- They provide empirical evidence across all their settings that reconstruction quality can be inversely correlated with feature usefulness, cautioning against using reconstruction quality as a primary measure of interpretability.
Main Findings
-
Toy model — better feature recovery despite worse reconstruction: With orbit size 4 and across orbit sizes 8, 16, and 32 in the appendix, Equivariant SAEs achieved lower reconstruction quality than baselines but outperformed them on feature recovery (mean cosine similarity to ground truth features via Hungarian matching) and feature detection (top-m latent logistic probes), with the gap widening as the number of ground truth features and interference increased (N = 8 undercomplete, 16 complete, 18 mildly overcomplete, 32 strongly overcomplete, with D = 16).
-
Activation transformations are largely linear: A learned transformation explains more than 97% of the variance in activations resulting from input rotations across the CNN, MLP, and ViT base models on the Shapes dataset. On GalaxyMNIST the trained transformation reached R² of 0.94–0.95 versus an identity (M = I) baseline of 0.91–0.92, and on MLL23 it reached 0.88–0.89 versus 0.78–0.81.
-
Shapes dataset — best probing across all setups: Probes trained on Equivariant SAE outputs achieved the best F1 scores across TopK values and truncation lengths under both C4 (90° rotations) and D4 (C4 plus vertical flips) symmetries, for both MLP and ViT activations. On ViT activations under D4, non-equivariant SAE probes struggled while Equivariant SAE probes reached close to 50% accuracy.
-
Reconstruction advantage depends on the base model: Equivariant SAEs had lower reconstruction error than narrow baselines on the MLP and CNN autoencoders, in both normalized reconstruction MSE and base model MSE when substituted into the model, but lagged behind on the ViT encoder despite better probing.
-
Symmetry benefit scales with base-model non-invariance: The CNN autoencoder was less invariant to C4/D4 transformations than the ViT encoder, and Equivariant SAEs provided larger probing gains over baselines on the CNN autoencoder, implying the benefits are greater when base activations vary more under input rotations.
-
Galaxy images (GalaxyMNIST, 10,000 images, 4 invariant classes): Equivariant SAEs gave the most accurate probes across sparsity budgets (TopK) and probe feature counts, reaching the accuracy of probes trained over base model activations as the number of features increased. The Canonical reconstruction and Output invariance objectives were considerably more accurate than the Latent invariance objective here.
-
Blood cell images (MLL23, 41,906 images, 18 cell types): Equivariant SAEs again gave the most accurate probes across TopK and probe feature counts. Unlike GalaxyMNIST, the Latent invariance objective performed best and the Canonical and Output objectives were less effective.
-
Reconstruction and usefulness can be inversely correlated: Vanilla SAEs had the lowest reconstruction errors while Equivariant SAEs generally performed worse, and on GalaxyMNIST the worst-performing Equivariant SAE for probing (the Latent invariance objective) gave the best reconstructions among the three Equivariant SAEs.
Methodology in Plain English
The authors start from the idea that a network's activations are sparse sums of feature directions. They ask what happens when the input is transformed by a symmetry such as a rotation: either the activations do not change (the model is invariant), or they change in a structured way. They hypothesize that for many transformations of interest, this change can be approximated by multiplying the activations by a matrix. Instead of making the SAE learn a separate feature direction for every orientation of the same concept, they train the SAE to map every rotated version of an input back to the activation of a canonical (unrotated) input, and separately learn matrices that transform the canonical reconstruction into the rotated activation.
Concretely, they handle groups that are products of two cyclic subgroups, requiring at most two learned matrices. The SAE encoder is a two-layer ReLU MLP with a TopK output activation, and they test three invariance losses: Canonical (reconstruct the canonical sample from each orbit), Latent (match encoder latents for canonical and transformed inputs), and Output (match the SAE reconstructions for both). The SAE and the transformation matrices are trained in two independent steps with different objectives; the matrices are initialized as identity matrices and optimized with Adam.
They compare against five baselines: vanilla and Archetypal TopK SAEs in narrow and wide versions (wide has 4× the latents, corresponding to allocating a separate latent per semantic feature per orientation), and group crosscoders, all trained on augmented data. Because evaluating interpretability is an open problem and auto-interp relies on language models that are themselves uninterpretable, they evaluate usefulness pragmatically: they measure feature recovery, detection, and reconstruction on a toy model with known ground truth, and on real activations they train XGBoost or logistic probing classifiers on SAE latents or reconstructions to predict known properties of the inputs. All results are averaged over three random seeds with standard errors shown.
Why This Matters
The paper makes a case that interpretability tools should inherit the same symmetry priors we already use to build more data-efficient scientific models, and that the field's default metric — reconstruction error — can point in the wrong direction. If features that reconstruct well are not the features that support downstream tasks, then SAE evaluation and SAE design both need rethinking, especially in scientific settings where the ground truth structure (rotation, orientation, physical symmetry) is known.
Real-world applications implied by the paper's settings and citations:
- Analysis of scientific models in domains such as proteins, cell images, and molecules, where the authors note symmetry aspects are currently overlooked.
- Astronomy: interpreting models trained on galaxy images, such as the ConvNeXT encoder pretrained on 1 million galaxy images and fine-tuned on 170,000.
- Computational pathology and hematology: interpreting DinoBloom-style vision transformers trained on over 380,000 single blood cell images from 13 hematology datasets.
- Any pipeline where an interpretability method must be audited by whether its features help downstream classification rather than by how faithfully they reconstruct activations.
Industry relevance: Teams deploying SAEs for model auditing, safety, or scientific discovery would need to account for data symmetries; the result that allocating 4× latents (the "wide" baseline setting) does not fix the problem suggests that simply scaling SAE capacity is not a substitute for correct priors. The two-step training design also allows the SAE and the transformation matrices to be trained on different datasets, which the authors note could support domain-specific SAEs.
Future Directions
- Generalizing the approach to larger or continuous groups, beyond the products of two cyclic subgroups handled here.
- Leveraging the learned transformation matrix for symmetry-aware feature labeling.
- Investigating how Equivariant SAE features interact across layers using methods such as Circuit Tracing.
- Resolving the open problem of SAE evaluation, given the paper's finding that reconstruction quality can be inversely correlated with feature usefulness; the authors explicitly state that probing over SAE latents or outputs is not claimed to be the optimal approach, since probing over base model activations often gives better results.
Target Audience
Mechanistic interpretability researchers working on SAEs and feature evaluation; equivariant and geometric deep learning researchers interested in how symmetry priors transfer from model design to model interpretation; and scientific machine learning practitioners in domains such as astronomy, biology, and medical imaging who need interpretable models over data with known symmetries. The paper is also relevant to anyone who currently uses reconstruction MSE as the primary yardstick for comparing interpretability methods.
Authors’ abstract
Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity. In particular, their activations entangle many concepts into fewer dimensions, a phenomenon known as superposition. Mechanistic interpretability methods such as sparse autoencoders (SAEs) can disentangle these dense activations into sparse sums of interpretable features, but SAEs suffer from unidentifiability: different explanations can fit the data equally well without necessarily being more interpretable or faithful to the underlying model. We show that this problem is exacerbated by data symmetries such as rotations that are prevalent in scientific domains. We extend the Linear Representation Hypothesis, the theory behind SAEs, to account for symmetries and show on synthetic as well as real-world scientific datasets and models that the resulting Equivariant SAEs can (1) avoid the pitfalls of existing SAEs on symmetric data and (2) discover features more useful for downstream tasks despite worse reconstructions. Our results show that incorporating the correct priors in SAEs can significantly improve their usefulness while highlighting that reconstruction quality can be inversely correlated with feature usefulness under symmetries, cautioning against its use as a key measure of interpretability. Code: https://github.com/ege-erdogan/equivariant-sae