Skip to content
AI.info

Research

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

Overview Research area: Natural Language Processing / mechanistic interpretability of transformer language models, specifically lens-based methods for reading intermediate hidden states. Technical lev

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
arXiv
2609.01936
Published
2026-09-01
Authors
Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane

AI summary

Overview

Research area: Natural Language Processing / mechanistic interpretability of transformer language models, specifically lens-based methods for reading intermediate hidden states.

Technical level: Advanced. The paper assumes familiarity with unembedding matrices (LM heads), the logit lens and its fitted variants, sparse autoencoders, and logit-margin analysis.

Scope: The paper introduces Sparse Readout Prism (SRP), which factorizes a model's unembedding matrix from its weights alone and expresses any token logit or logit difference as a sum of signed contributions from sparse readout features, then uses that basis to show that fitted lenses report the language of their fitting corpus while the dominant readout feature stays fixed.

What This Paper Is About

Lens methods such as the logit lens, tuned lens, and Jacobian lens decode a model's intermediate hidden states into tokens to trace how a next-token prediction develops across layers. But a lens reading reflects two things at once: the hidden state and the readout (the unembedding matrix) used to decode it. Because many lenses are fit on a corpus, two lenses differing only in their fitting corpus can report different tokens for the same hidden states, a dependence the authors call corpus conditionality. SRP changes the unit of analysis from token identities to readout features, building one basis per model from the readout weights alone—no corpus—so that lens readings can be compared across tokens, contexts, layers, and lenses.

Key Contributions

  1. SRP, a corpus-free decomposition method. SRP applies dictionary learning directly to the rows of the LM head, writing each row as sparse coefficients over a shared overcomplete basis plus a residual. Each selected score (a token logit or a margin between two tokens) then splits into one signed contribution per feature plus a residual, and the offset, contributions, and residual sum to the score exactly by construction.

  2. Fidelity evidence for the decomposition. Replacing the original readout by SRP's sparse approximation reconstructs 8.9 to 17.3 percentage points more of the tested logit differences than the strongest of six baselines built on geometric relations among readout rows, and ablating a feature shifts a logit difference in proportion to its SRP contribution.

  3. The corpus conditionality finding. Two Jacobian lenses fit on English and Chinese C4 prompts report tokens in different scripts for the same frozen states on 39 of 80 prompts, while the same SRP feature dominates under both lenses on 77 of 80 prompts (96%, 95% CI [0.90, 0.99]).

  4. A warning and a control. The results caution against treating token rankings as direct evidence about intermediate computation, and supply a reference fixed in the weights before any reading is taken.

Main Findings

  • Corpus conditionality is real for fitted lenses. Fitting two Jacobian lenses on Qwen3.5-9B that differ only in their fitting corpus—100 English versus 100 Chinese C4 prompts—produces different reported languages for identical hidden states. On factual recall prompts the lenses report different scripts on 9 of 10 cases. On the antonym prompt from Gurnee et al. (2026), the English-fitted lens reports large at layers 24 and 26 and big at layer 29, while the Chinese-fitted lens reports 大的 at layers 24 and 26 and 大 at layer 29.

  • The dominant readout feature is stable even when tokens differ. The same SRP feature dominates under both lenses on 77 of 80 prompts (96%, 95% CI [0.90, 0.99]). In the paper's Figure 4 example, feature f23180 accounts for 75% of the feature sum for the English-fitted lens and 63% for the Chinese-fitted lens, with every other feature contributing at most +1.7 logits. Null comparisons produce no matches: 0/159 for unrelated tokens and 0/240 after shuffling prompts.

  • Stability extends across lens constructions, scripts, and corpus sizes. With ridge translators related to the tuned lens, the same feature dominates on 67/80 prompts (CI [0.74, 0.90]). Holding the English fitting corpus fixed, the Jacobian lens and ridge translator agree on the dominant feature for 76/80 prompts. For English–German, which share the Latin script, the lenses report different top-1 tokens on 73/90 prompts while the same feature dominates on 84/90. Tripling each fitting corpus from 100 to 300 prompts leaves feature agreement at 77/80 (English–Chinese) and raises it from 84/90 to 85/90 (English–German), while token disagreement moves in opposite directions across the two pairs.

  • Replacement fidelity splits by model family. Replacing the readout with SRP's reconstructed LM head preserves the held-out argmax on 0.89–0.90 of decoded states for the Qwen3.5 and Ministral readouts and on 0.75–0.76 for the two readouts distilled for reasoning. Median KL between original and reconstructed distributions is 0.09–0.14 bits for Qwen3.5 and Ministral and 0.43–0.49 bits for the distilled-for-reasoning readouts.

  • SRP beats six row-geometry baselines. Coverage reaches 0.72–0.81 across the six softcap-free readouts (Q-0.8B 0.754, Q-2B 0.780, Q-9B 0.812, Min-8B 0.777, R1-Q-7B 0.716, R1-L-8B 0.751), against best-alternative values of 0.665, 0.625, 0.639, 0.648, 0.579, and 0.624, with non-overlapping intervals. A k-means dictionary matched to SRP's own width (D = 65,536) and sparsity (k = 256) trails SRP by 18 points of coverage on Qwen3.5-2B. Shuffled codes and random support leave coverage below 0.10 on the three Qwen3.5 and both Gemma-4 readouts.

  • Contributions predict ablation effects. Removing a feature's direction from the decoded state shifts the margin in line with the predicted contribution at r² = 0.83–0.93 on all six softcap-free readouts (Q-0.8B 0.886, Q-2B 0.834, Q-9B 0.829, Min-8B 0.926, R1-Q-7B 0.918, R1-L-8B 0.934), against at most 0.06 for a random direction. Slopes are 0.96–1.13 on Ministral and the two distilled-for-reasoning readouts, while on the three Qwen3.5 readouts the margin moves further than predicted (slope 1.34–1.69).

  • Decompositions survive retraining. Three dictionaries per width differing only in seed (32× and 16× on Qwen3.5-2B) all replace the readout equally well, with held-out top-1 agreement of 0.805–0.848. Feature directions themselves do not recur across seeds (median cosine of about 0.33 between a direction and the nearest direction in another seed's dictionary), yet the tokens a dictionary groups with a given token reproduce at a mean Jaccard of 0.21–0.24 versus 0.014–0.018 for unrelated contrasts—13–15× the overlap—and 87–91% of the features find a counterpart in the other seed's dictionary.

  • Context decides which features carry a token. On Qwen3.5-2B, the margin favoring bug is carried by mosquito, bee, and beetle features in an insect context (ρ₀ = 0.185) and by defect, crash, and debug features in a software context (ρ₀ = 0.003), with no insect-sense feature among the leading terms in the software case.

  • Features discriminate word senses better than chance. On 20 ambiguous words from CoarseWSD-20, balanced accuracy is 0.90, 0.92, and 0.80 on Qwen3.5-2B, Qwen3.5-9B, and R1-Llama-8B, where shuffling sense labels gives a null near 0.41. A nearest centroid probe on the full decoded state reaches 0.96, 0.97, and 0.92, so the full state is the stronger predictor and eight features give a readable summary of it.

Methodology in Plain English

The setup. In a transformer, each output logit is a dot product between the final state presented to the LM head and one row of the unembedding matrix W_U. SRP takes that matrix on its own—no prompts, no corpus—and learns a sparse dictionary over its rows.

Fitting the dictionary. The authors fit a TopK sparse autoencoder to the unembedding rows, fixing k = 256 at 32× expansion for their reported operating points. TopK means the encoder keeps the k largest pre-activations in each row, so the sparsity budget is set directly. Rows are centered and their norms normalized before fitting, and both steps are inverted afterward so every term lands back in original logit units. Each row becomes a fitted row (a shared offset plus sparse coefficients over the feature directions) plus a residual.

Turning a score into feature contributions. Any score the authors study—a token's logit, the margin between two tokens, or a comparison involving an averaged token set—is written as a coefficient vector over unembedding rows. Substituting the row decomposition and rearranging gives the score as a shared offset term, one signed contribution per readout feature (the feature's coefficient in the selected direction times how strongly the decoded state aligns with that feature direction), and a residual. For contrasts, which are scores whose coefficients sum to zero, the offset term vanishes and the score reduces to signed contributions plus residual.

Checking the answers. Three diagnostics accompany each decomposition: replacement fidelity (does substituting the reconstructed LM head preserve held-out argmaxes), relative reconstruction error (residual divided by score, reported at a floored setting ρ_0.5 for aggregates and the stricter ρ_0 for case studies and baselines), and sign agreement (does the reconstruction carry the same sign as the exact score). Coverage combines sign agreement with ρ₀ below one half.

The evidence suite. The authors span Qwen3.5, Gemma-4, Ministral, and Qwen/Llama readouts distilled for reasoning at 0.8B–9B scale. They fit one readout SAE per model to its final LM head and evaluate on the decoded states of 10,000 C4 continuations that never enter the fit, using banks of roughly 1,350 scores per model drawn from about 850 contrasts.

The corpus-conditionality experiment. They hold an intermediate hidden state fixed and vary only the corpus used to fit the lens, then decompose the score assigned to the token each lens reports and record the feature with the largest absolute contribution. A prompt counts as agreement when the same feature dominates under both lenses in at least two of layers 21, 24, 26, and 29. Two null comparisons estimate chance agreement: evaluating an unrelated token under the same lens, and shuffling prompts before pairing the lenses.

Why This Matters

Lens readings are frequently read as evidence about a model's internal computation—for example, treating English intermediate tokens as evidence that a model reasons in English internally. The paper shows this inference is underdetermined: a fitted lens's report can change with the corpus it was fit on, while the readout structure underneath stays put. Because SRP's basis is built before any reading is taken and uses no corpus, it gives that inference the control it needs, and it also gives interpretability work a unit finer than token identity, which cannot separate senses of a single token or recognize structure shared across tokens.

Real-world applications:

  • Safety audits. The authors propose using SRP wherever a question admits a concrete readout score, defining group contrasts for refusal, source verification, or a specified lexical family, then comparing feature contributions across paraphrases, languages, adversarial prompts, and checkpoints.
  • Constrained model edits. The feature terms supply candidate directions for readout-side edits. At matched next-token KL those directions beat the mean and leading principal components of the rows they were discovered from

Authors’ abstract

A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. Many lenses are fit on a corpus, and we show that two lenses differing only in their fitting corpus can report different tokens for the same hidden states. We call this dependence corpus conditionality. To examine readout structure independently of the fitting corpus, we introduce Sparse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features. This reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses. Replacing the original readout with SRP's sparse approximation reconstructs 8.9-17.3 percentage points more of the tested logit differences than the strongest of six baselines built on geometric relations among readout rows. Ablating features shifts logit differences in proportion to their SRP contributions. Although token readings vary with the fitting corpus, the dominant readout feature remains stable. Because SRP uses no corpus in its construction, it provides a control independent of the fitting corpus for lens analyses.

Read the original paper