Research
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
Summary: Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders Overview Research area: Mechanistic interpretability of large language models, specifically feature-discovery

- arXiv
- 2510.22332
- Published
- 2025-10-25
- Authors
- Mengyu Ye, Jun Suzuki, Tatsuro Inaba, Tatsuki Kuribayashi
AI summary
Summary: Transformer Key-Value Memories Are Nearly as Interpretable as Sparse AutoencodersOverview
Research area: Mechanistic interpretability of large language models, specifically feature-discovery methods for Transformer feed-forward (FF) layers.
Technical level: Intermediate. The paper assumes familiarity with the Transformer architecture, feed-forward layers, sparse autoencoders (SAEs), and standard interpretability evaluation terminology, but its arguments are conceptual rather than mathematically heavy.
Scope in one sentence: The paper systematically compares the interpretability of features already stored in FF layer key-value memories against features learned by proxy modules (SAEs and Transcoders), using the SAEBench suite plus manual human evaluation.
Note on the supplied content: The paper text provided here is truncated mid-table in Section 4. The full numerical results table is not reproduced, so this summary reports the qualitative findings and the specific figures that do appear in the available text.
What This Paper Is About
Recent LLM interpretability research has been dominated by a proxy-based approach: train an external module such as a sparse autoencoder to decompose neuron activations into simpler, supposedly more interpretable features, then evaluate those learned features. This raises a question the authors argue has rarely been tested systematically — do features learned by a proxy actually have better properties than the features already represented in the original model parameters?
The paper addresses this by treating feed-forward layers as key-value memories, where key activations act as features and value vectors act as feature vectors, and then evaluating those "organic" features with the same modern benchmarks used to judge SAEs.
Key Contributions
-
A direct, benchmark-based comparison of FF key-value features against proxy-discovered features. The authors evaluate FF-KV features, SAEs, and Transcoders with SAEBench metrics, reporting that SAE and FF features fall into a similar range of interpretability across all eight SAEBench metrics.
-
Three FF-KV variants that mirror SAE design choices. The paper introduces vanilla FF-KV, TopK FF-KV (a top-k activation function applied to the key vector, keeping only the k largest activations and zeroing the rest), and Normalized (TopK) FF-KV (normalizing each row of the value matrix and reweighting activations by the discounted vector norms), including an implementation compatible with SwiGLU gating.
-
Human manual evaluation of feature quality. Independent of the automated metrics, the authors manually inspect features and report that conceptual features can be found with almost equal ease from both FF-KV features and SAEs.
-
A faithfulness analysis treating FF-KV features as gold. Using Transcoder (TC) as the closest proxy counterpart to FF-KV, the authors measure how much the proxy's feature set overlaps with the original FF module's feature set, and find that the majority of TC features have no similar counterpart in the original FF.
Main Findings
-
Comparable interpretability across the board: Evaluation scores fell into a similar range in all eight metrics in SAEBench for FF-KV and SAE features. The authors characterize SAE improvements as observable but minimal in some aspects.
-
Parallel inter-metric tendencies: The two approaches behave similarly relative to each other across metrics — for example, causal intervention scores are poorer than feature disentanglement scores in both methods.
-
FF-KV sometimes wins: Features in the original FF layers tend to avoid feature overlapping, leading to better absorption scores (interpreted as less redundancy) than SAEs. The authors note this is a surprising result in favor of vanilla FFs.
-
Divergent feature sets: Features discovered in SAEs and in FFs diverge, raising questions about the advantage of SAEs from both feature-quality and faithfulness perspectives.
-
Proxy features are largely unmatched in the original model: Most Transcoder features do not have similar counterparts in the original FF module. This aligns with existing reports that SAEs can interpret even random Transformers, suggesting the proxy may hallucinate new features rather than translate the workings of the original FF module.
-
Perfect reconstruction by construction: The FF-KV methods, which are not proxies, automatically receive a perfect explained variance score (= 1) because they are the original module rather than an approximation of it.
-
Manual evaluation supports the automated results: In the authors' human assessment, conceptual features are found with almost equal ease from FF-KV features and SAEs.
-
Metrics instability caveat: The authors report that the Spurious Correlation Removal (SCR) and Targeted Probe Perturbation (TPP) scores are highly unstable and should be treated as supplementary. For those metrics they consider top-K activations with K = 20, following existing studies.
-
Gemma-2 2B is the model named in the available results table, evaluated with an SAE and a Transcoder configuration; the remaining table rows and their numeric values are not present in the supplied content, so per-metric figures for the full comparison are not reported here.
Methodology in Plain English
The authors start from the observation that a Transformer feed-forward layer already performs a decomposition that looks structurally like what an SAE does: it projects the input into a higher-dimensional space (with d_FF typically set to 4 × d_model), applies a non-linearity, and then combines a set of value vectors weighted by the resulting activations. Instead of training a new module to explain those activations, they simply treat the FF layer's own key activations as features and its value vectors as feature vectors, and run the standard evaluation machinery on them.
To make the comparison as fair as possible, they adapt SAE-style design choices onto the FF layer: sparsifying activations with a top-k function, and normalizing value vectors so that features with large norms do not have their activations underestimated. They also implement the methods on top of SwiGLU, the gating function used in modern language models, following prior work showing SwiGLU's compatibility with FF-KV analysis.
They then evaluate SAEs, Transcoders, and the FF-KV variants with the SAEBench metrics, which include Feature Alive Rate, Explained Variance, Absorption Score, Sparse Probing, Auto-Interpretation Score, Spurious Correlation Removal (using the SHIFT data), its extension TPP, and RAVEL, which is reported as an Isolation score and a Causality score. They complement the automated evaluation with manual inspection of features by human annotators, and finally measure faithfulness by checking how many Transcoder features correspond to features actually present in the original FF module.
Why This Matters
Impact on research: The paper argues that FF key-value parameters serve as a strong baseline for modern interpretability research, and that the theoretical advantage of SAEs is not observed empirically under the current evaluation scheme. It also questions the faithfulness of proxy-discovered features, encouraging the community to be explicit about what a proxy is actually explaining.
Real-world applications:
- Model auditing and safety review: Teams inspecting internal model behavior can examine FF key and value parameters directly rather than training and interpreting a separate proxy module, and can use FF features as a reference point when validating whether a proxy's explanation is faithful.
- Behavior steering and control: Because FF key-value features are interpreted as knowledge-retrieval units whose activation and value vectors can be manipulated, FF-KV features are candidates for the same kind of activation-level interventions used with SAE features.
- Interpretability tooling for smaller organizations: The paper notes that proxy-based pipelines add computation cost to interpret a model; direct FF analysis avoids training a separate module, lowering the resource barrier to feature-level analysis.
- Evaluation of third-party interpretability products: When an interpretability method or tool claims to recover the "true" features of a model, FF-KV analysis provides a grounding point and a simpler baseline for checking that claim.
Industry relevance: The findings matter to anyone building or relying on feature-discovery infrastructure for LLMs — safety and alignment teams, evaluation groups, and researchers who need to justify the extra cost and complexity of training SAEs or Transcoders over directly reading out what the model's own parameters already encode.
Future Directions
- Faithfulness evaluation with FF-KV features as ground truth: The authors explicitly call for further research on the faithfulness of learned proxy features, using FF-KV features as grounding points rather than treating the proxy's own output as the final word.
- Determining what proxies actually learn: Given that most Transcoder features lack counterparts in the original FF module, the field needs a clearer account of whether proxy features are useful generalizations, hallucinations, or useful abstractions that do not map onto any single model feature.
- Stabilizing the disentanglement metrics: The paper reports that SCR and TPP scores are highly unstable and investigates different values of K, which leaves open the question of how to build reliable measures of feature separation.
- Extending the comparison beyond the evaluated settings: The supplied content names Gemma-2 2B and describes FF-KV variants, SAEs, and Transcoders, but does not report whether the same conclusion holds for attention layers, other model sizes, or other architectures — a natural next step.
Target Audience
This paper is most useful for mechanistic interpretability researchers working on feature discovery, engineers and safety practitioners who build or depend on SAE- and Transcoder-based analysis pipelines, and graduate students entering the interpretability field who want to understand why a simpler baseline might be competitive with a much more expensive approach. A reader should be comfortable with Transformer feed-forward layers, the key-value memory framing, and the basic vocabulary of sparse autoencoder training and evaluation.
Authors’ abstract
Recent interpretability work on large language models (LLMs) has been increasingly dominated by a feature-discovery approach with the help of proxy modules. Then, the quality of features learned by, e.g., sparse auto-encoders (SAEs), is evaluated. This paradigm naturally raises a critical question: do such learned features have better properties than those already represented within the original model parameters, and unfortunately, only a few studies have made such comparisons systematically so far. In this work, we revisit the interpretability of feature vectors stored in feed-forward (FF) layers, given the perspective of FF as key-value memories, with modern interpretability benchmarks. Our extensive evaluation revealed that SAE and FFs exhibits a similar range of interpretability, although SAEs displayed an observable but minimal improvement in some aspects. Furthermore, in certain aspects, surprisingly, even vanilla FFs yielded better interpretability than the SAEs, and features discovered in SAEs and FFs diverged. These bring questions about the advantage of SAEs from both perspectives of feature quality and faithfulness, compared to directly interpreting FF feature vectors, and FF key-value parameters serve as a strong baseline in modern interpretability research.