Skip to content
AI.info

Research

Latent Introspection: Models Can Detect Prior Concept Injections

Overview Research area: AI interpretability and model introspection — specifically whether language models can access and report on their own prior internal states. Technical level: Intermediate. The

Latent Introspection: Models Can Detect Prior Concept Injections
arXiv
2602.20031
Published
2026-02-23
Authors
Theia Pearson-Vogel, Martin Vanek, Raymond Douglas, Jan Kulveit

AI summary

Overview

Research area: AI interpretability and model introspection — specifically whether language models can access and report on their own prior internal states.

Technical level: Intermediate. The paper assumes familiarity with transformer internals (residual streams, KV caches, logit lens, steering vectors) but explains its design choices clearly.

Scope: An empirical study showing that a single open-weight 32B model (Qwen2.5-Coder-32B-Instruct) carries detectable information about concepts injected into its earlier context, even though its sampled outputs deny the injection.

What This Paper Is About

The authors ask whether a language model can tell when a concept has been artificially injected into its own earlier internal states. They adapt a paradigm introduced by Lindsey (2025) on proprietary Anthropic models and apply it to an open-weight model, then probe more thoroughly than output sampling alone would allow. The central puzzle is that the model's actual text output says "no," while analysis of its middle layers shows it has clear information that the injection occurred and which concept was involved.

Key Contributions

  1. Demonstrating that an open-weight 32 billion parameter model (Qwen2.5-Coder-32B-Instruct) can detect prior concept injections, extending Lindsey (2025) to a setting the research community can reproduce and build on.
  2. Showing that detection capacity can be too weak for standard sampling-based evaluation yet still visible through analysis of intermediary layers via logit lens.
  3. Showing that prompting can elicit balanced accuracy of up to 84.0%.
  4. Recovering injected concepts with up to 1.36 bits of mutual information, and showing this capacity correlates with detection sensitivity across prompts (Pearson r = 0.68, p = 0.004).

Main Findings

  • Detection is real but hidden in outputs: In the baseline configuration, the model's most likely next token stays "no" regardless of whether an injection happened. Logit lens analysis reveals the detection signal in middle layers (50–60) before it attenuates in the final layers.

  • Prompting dramatically changes sensitivity: Asking plainly whether a concept was injected moves P("Yes") from 0.2% (false positives) to 0.3% (true positives). Adding detail about how a model might access prior hidden states raises sensitivity to 39.9% while increasing false positives by only 0.6%.

  • The effect is specific to introspection: Always-yes and always-no control questions shift by less than 0.2% under injection; genuinely ambiguous questions shift by at most 8.0%; "confusing" questions (unambiguous answers the model often gets wrong) shift by up to 25.8%.

  • Which concept can be recovered: Across nine injected concepts (cats, bread, love, fear, death, truth, creativity, programming, music), the model identified programming at 92%, death at 75%, and truth at 71%. Two of the nine (cats and bread) were not identified correctly. Mutual information reached 1.36 bits, or 43% of the theoretical maximum of log₂(9) ≈ 3.17 bits. The abstract additionally reports mutual information rising from 0.61 bits to 1.05 bits.

  • Signals emerge late and are suppressed late: Although the steering vector is applied at layers 21–42, introspection signals do not emerge until roughly layer 48, peak around layers 58–62 (with P("yes") approaching 100% under injection), and then drop sharply in the final 2–3 layers. Concept-identification MI peaks at layers 61–62.

  • Vague framing beats accurate mechanism descriptions: The Vague Mechanism framing achieves 68–84% balanced accuracy across all info conditions; Accurate Mechanism reaches only 42–70% and needs the info document to exceed chance. The Wrong Mechanism framing, which inaccurately describes the intervention as full finetuning, performs similarly to the accurate one. The Poetic Document condition actually decreases P("yes") under injection from a baseline of 47%.

  • Information scaffolding helps selectively: The pro-introspection document improves detection by 20–32 percentage points for the accurate and inaccurate mechanism framings but has minimal effect on the vague framing. The vague framing performs best with lorem ipsum filler (84%).

  • Detection and identification move together: Conditions that improve concept identification also improve detection sensitivity (Pearson r = 0.68, p = 0.004), suggesting a shared underlying capacity rather than a task-specific artifact.

  • Replication on larger models: Llama 3.3 70B Instruct and Qwen 2.5 72B Instruct both show introspection signals. Qwen-72B shows greater accuracy with the accurate mechanism framing plus the pro-introspection document. Llama-70B shows an inverted info document effect: providing introspection-relevant information decreases accuracy (75.5% vs. 38.0%). All three models show final-layer attenuation, with peak MI layers scaling with model depth.

Methodology in Plain English

The researchers needed a way to check whether a model truly reads its own past internal states, rather than just reacting to an ongoing perturbation or guessing from what it said earlier.

They start by training "steering vectors" — directions in activation space — for each of nine concepts using the repeng library with PCA-based extraction. A vector is built by contrasting activations from prompts that mention the concept against neutral prompts. They verify that these vectors work: applying the "cat" vector during generation makes the model spontaneously discuss cats.

The key trick is when they apply the vector. They apply it only while generating the key-value cache for an initial conversation turn, over the middle layers of the model (layers 21–42 of 64). Then they remove the steering vector entirely. Only afterward do they append a second user message asking whether a concept was injected, along with a preset assistant prefix ("The answer is"), and read off the probability of "yes" versus "no."

Because the first turn is already fixed and the vector is gone by query time, the model cannot be reacting to ongoing perturbation or inferring the injection from its own steered language. Any detection must come from information cached in earlier representations.

They crossed four framings (Accurate Mechanism, Wrong Mechanism, Vague Mechanism, Poetic No Mechanism) with four info documents (none, pro-introspection text, matched lorem ipsum filler, poetic text) for 16 conditions. For concept identification, they present a numbered list of nine concepts plus a "no injection" baseline, extract logits over digit tokens using logit lens, and average over five shuffled list orderings per trial. They also use four categories of control questions (16 always-no, 8 always-yes, 6 varied-baseline, 4 confusing) to check that injection does not simply bias responses. Results use nine concepts and ten random seeds for steering-vector training unless stated otherwise.

Why This Matters

Impact on research: The paper argues that behavioral evaluation alone — reading sampled outputs — may systematically underestimate what a model knows about itself. The open-weight setting and released code let other researchers reproduce and extend the result, unlike prior work on proprietary models.

Real-world applications:

  • Safety evaluations: assessment pipelines that rely solely on what a model says could miss capabilities it actually possesses.
  • Alignment training: if honesty training penalizes claims of introspective access, it might incentivize models to under-report real capabilities.
  • Interpretability tooling: logit lens and mutual-information analysis offer a template for surfacing latent knowledge that never reaches the output layer.
  • Latent reasoning research: access to prior internal states may be a precursor to latent reasoning, making detection relevant to how models plan or deliberate internally.

Industry relevance: The finding that prompt wording swings balanced accuracy from near-chance (50.1%) to 84% means that whether a capability shows up in an evaluation may depend heavily on how the question is phrased. Teams deploying or auditing models may need probing methods beyond conversation, and the paper's replication showing an inverted info-document effect between Llama-70B and Qwen-32B indicates that results do not transfer straightforwardly across model families.

Future Directions

  • Why is introspection suppressed? The authors propose three unexplored hypotheses for the late-layer attenuation: post-training effects (RLHF penalizing self-awareness claims), pretraining factors, and distribution shift pushing introspection queries out of distribution. They suggest comparing base and instruction-tuned variants to isolate post-training effects.
  • Why does vague framing outperform accurate description? Two competing explanations remain untested: that accurate mechanism descriptions trigger learned denial responses, or that "what seems salient right now" is more naturally represented than "what was done to my activations."
  • Mechanistic account. The paper observes where signals emerge and attenuate but does not identify the responsible circuits or causally intervene on them.
  • Prompt sensitivity and cross-model generalization. The large gap between prompts is unexplained, and the authors note that the models they tested respond very differently to the same prompt manipulations, leaving open how far these results generalize.

Target Audience

AI interpretability and alignment researchers, especially those working on model introspection, latent knowledge, and self-report reliability. Safety evaluators and policy-adjacent readers interested in whether behavioral testing can miss capabilities will also benefit, as will engineers working with open-weight models who want a reproducible method for probing internal representations. Readers without background in transformer internals will need to consult the cited work on logit lens, steering vectors, and KV caches to follow the methods in detail.

Authors’ abstract

We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% -> 39.2%) with only a 0.6% increase in false positives. Also, mutual information between nine injected and recovered concepts rises from 0.62 bits to 1.05 bits, ruling out generic noise explanations. Our results demonstrate models can have a surprising capacity for introspection and steering awareness that is easy to overlook, with consequences for latent reasoning and safety.

Read the original paper