Research
Selecting The Most Informative Tokens in Natural Language Autoencoders
Overview Research area: Interpretability and auditing of large language models, specifically Natural Language Autoencoders (NLAs) that translate a model's internal activations into readable text, comb

- arXiv
- 2609.37040
- Published
- 2026-09-29
- Authors
- Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez, Amon Rapp, Luigi Di Caro, Peter Schneider-Kamp, Lukas Galke Poech
AI summary
Overview
Research area: Interpretability and auditing of large language models, specifically Natural Language Autoencoders (NLAs) that translate a model's internal activations into readable text, combined with token-position selection for efficient auditing.
Technical level: Advanced. The paper assumes familiarity with transformer internals (residual stream, attention patterns, next-token distributions), reinforcement-learned verbalizers, AUROC-based ranking evaluation, and fine-tuning adapters.
Scope: The paper measures which token positions an auditor should inspect when generating NLAs, comparing thirteen computational signals against a ranker trained only on chat structure, across 4,705,657 explanations from 4 models and 4 datasets.
What This Paper Is About
Generating an NLA explanation is expensive: the verbalizers released by Fraser-Taliente et al. (2026) generate 130 tokens on average per activation (maximum 150), while a transcript under audit contains hundreds to thousands of positions. Previous work picked positions by convention, such as reading only the final position before an answer or using eight evenly spaced positions per transcript. This paper asks whether a score computable before any explanation is generated can rank the positions whose explanations will actually describe the threat under audit, and whether the rendered transcript's chat structure alone is enough for that ranking.
Key Contributions
- The first systematic study of NLA position selection across 4,705,657 explanations, linking selection scores to explanation relevance at every token position.
- A comparison of 247 candidate scores (13 signals plus 234 two-signal ensembles), showing that activation-vector signals are strongest among individual signals in 11 of 14 dataset-model combinations, and that combining signals improves ranking on held-out transcripts in all fourteen combinations.
- Evidence that chat structure exceeds the best individual signal in AUROC in 12 of 14 combinations, and that a 5% position budget retains 95.8% of the success rate from explaining every position on three datasets.
- A demonstration that pretrained verbalizers recover words models have learned to conceal through fine-tuning, without additional verbalizer training, with the same prompt producing explanations specific to each model's concealed word.
Main Findings
-
Activation signals lead, but only narrowly: The strongest signal in a cell is computed from the activation in 11 of 14 cells and from the attention pattern in the remaining 3; no predictive-distribution signal is strongest in any cell.
resid_jump_nlais strongest in five cells anddominant_massin three. The median direction-adjusted AUROC across signals and cells is 0.584, the middle half of values lies between 0.535 and 0.634, and the maximum AUROC in any cell is 0.796. -
Chat structure beats the forward pass: The
structurebaseline (segment, chat role, boundary ordinal, normalized index, requiring no forward pass) has higher pooled AUROC than the best single signal in 12 of 14 cells, and higher case-macro AUROC in 11 of 14. At a budget of one explanation,structuregives higher precision than the selected ensemble in 11 of 14 cells; at a budget of eight, in 12 of 14. -
Signal directions flip across datasets: Eight of the thirteen signals are in the minority direction in five or more of the fourteen cells.
entropyis the only signal with one direction across models within each dataset, and it still reverses between OpenPromptInjection and the taboo organisms. A signal applied with the wrong sign is less accurate than choosing positions at random. -
Ensembles help, modestly: All four datasets select two-signal ensembles. OpenPromptInjection selects
lookback_ratio+sink_drain(50/50; full/held-out 0.754/0.761, Δ = 0.104/0.106); Tensor Trust selectsresid_jump_nla+w(75/25; 0.682/0.567, Δ = 0.009/0.031); Liars' Bench selectsdominant_mass+head_disagreement(50/50; 0.722/0.541, Δ = 0.021/0.084); taboo organisms selectdominant_mass+norm_ratio(50/50; 0.750/0.775, Δ = 0.042/0.053). All held-out improvements have positive 95% confidence intervals. -
Explanations localize the threat: Positions inside the injected span in OpenPromptInjection receive an on-task explanation 13 to 66 times more often than positions outside it. On Liars' Bench the same comparison inside the graded reply gives 8 and 9 times for l70 and g27. Tensor Trust has a ratio below one across all four models.
-
On-task rates vary widely: The rate of on-task positions is 0.013 to 0.30 on OpenPromptInjection, the taboo organisms and Liars' Bench, and 0.68 to 0.86 on Tensor Trust.
-
A 5% budget is close to exhaustive: At a 5% position budget,
structuregives 0.995 on OpenPromptInjection, 0.958 on the taboo organisms and 1.000 on Tensor Trust; random choice gives 0.761, 0.594 and 1.000.structuregives 0.813 on Liars' Bench. Requiring three on-task explanations instead of one lowers the first three numbers to 0.785, 0.568 and 0.998. Median transcript lengths are 143, 74, 131 and 264 positions. -
Concealed words are recovered without retraining: Across the twelve taboo organisms (four base models each fine-tuned to conceal
moon,shiporsnow), an explanation states the secret word at 14% to 27% of positions. At positions of the four prompts that ask for the secret, the organism's own word is named at 12% to 25% of positions, while another organism's word is named at most 2.8% of positions, which text inversion cannot explain. -
Top-ranked single-signal positions fail at tiny budgets: On seven of the ten cells where on-task positions are sparse, the position the strongest single signal ranks first is on-task in no transcript. Mixing two signals moves that position off the spike: the selected pair gives 0.594, 0.625 and 0.771 on the three taboo organisms where the single signal gives 0.
-
Boundary preference differs by task: OpenPromptInjection ranks the boundary above the input in every model by +0.22 to +0.40; Tensor Trust by +0.06 to +0.38; the taboo organisms by −0.08 to +0.20 in three of four models; Liars' Bench ranks the boundary below the input in both models, by −0.40 and −0.32.
-
Controls behave as expected: The random score and the shuffled-label control are both 0.5 where the procedure is sound.
head_disagreementwas registered in advance as the primary signal for Liars' Bench and Tensor Trust; all other signals are exploratory under Benjamini-Hochberg control of the false discovery rate at q = 0.05.
Methodology in Plain English
The researchers use NLAs, which pair an Activation Verbalizer (AV) that writes a short paragraph describing an activation with an Activation Reconstructor (AR) that maps that paragraph back to an estimated activation, trained so that the reconstruction error |h − ĥ|₂² stays low.
For every token position of every transcript, they generate an explanation and ask a judge model (DeepSeek-V4-Flash at temperature zero) a fixed yes/no question specific to the dataset about whether the explanation concerns the threat being audited. A "yes" makes the explanation on-task.
They then test thirteen cheap scores, each computable in a single forward pass: four from the next-token predictive distribution (surprisal, entropy, varentropy, temporal KL), four from attention patterns (lookback_ratio, sink_drain, head_disagreement, and a combined w), and five from the activation vector (resid_jump, norm_ratio, peak_ratio, dominant_mass, resid_jump_nla). These are turned into ranks, oriented per dataset and model, and combined pairwise with weights α in {0.25, 0.50, 0.75}, yielding 247 candidates.
Selection happens either per dataset-model pair (model-best) or shared across all models of a dataset (dataset-shared). Evaluation uses five-fold validation over whole transcripts, with measures of precision at fixed budgets (one position, eight positions, 1%, 5%, 10% of a transcript), case-macro AUROC computed within transcripts, and pooled AUROC. Two baselines derived only from the transcript are trained as logistic regressions on four fifths of transcripts: position, and structure which adds chat role and segment.
Why This Matters
Impact on research: The paper reframes NLA position choice as a measurable selection problem and shows that the cheap path — the rendered transcript's chat structure — often beats activation- or attention-derived scores. It also shows that a verbalizer trained on one model can produce useful explanations for a fine-tuned copy with different weights, which challenges the assumption that a verbalizer only describes the model it was trained on.
Real-world applications:
- Live monitoring of deployed assistants, where explaining every token position is impractical but a 5% budget retains 95.8% of the success rate from explaining every position on three datasets.
- Prompt-injection detection in tool-using agents, using position rankings to concentrate inspection on injected instructions.
- Auditing fine-tuned or adapter-modified models for concealed behavior, where pretrained verbalizers named the secret word at 14% to 27% of positions.
- Budget planning for red-teaming pipelines, since precision did not vary across budgets of one position, 8 positions, 1% and 10%, letting an auditor scale cost with affordable compute.
Industry relevance: Explanation generation dominates audit cost because each explanation costs roughly 130 tokens of autoregressive decoding. If selection can be done with no forward pass, as the structure baseline shows, then monitoring systems can allocate explanation budgets by transcript layout rather than by running the audited model, and can evaluate the base rate from a small sample before choosing a selector.
Future Directions
- Replicate the on-task labels with a second evaluation model and human review, since the current study uses one evaluation model and therefore cannot test label reliability.
- Test generalization across verbalizer layers, as the study evaluates one fixed layer per model (layers 20, 32, 41 and 53 for q7, g12, g27 and l70).
- Extend to longer agent transcripts, which the authors list as untested given the relatively short transcripts used here.
- Move beyond selecting positions by task relevance toward checking the content of the explanation itself, which the conclusion frames as a first step.
Target Audience
Interpretability and safety researchers working on model auditing, activation verbalization, and monitoring of deployed language models; practitioners who need to budget explanation generation over long transcripts; and machine learning engineers evaluating fine-tuned or adapter-modified models for concealed behavior.
Authors’ abstract
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.