Research
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path Authors: Andres Saurez, Neha Sengar, Dongsoo Har (Korea Advanced Institute of Science and Technology, Daejeon 34051, South Korea)

- arXiv
- 2602.09784
- Published
- 2026-02-10
- Authors
- Andres Saurez, Neha Sengar, Dongsoo Har
AI summary
Circuit Fingerprints: How Answer Tokens Encode Their Geometrical PathAuthors: Andres Saurez, Neha Sengar, Dongsoo Har (Korea Advanced Institute of Science and Technology, Daejeon 34051, South Korea) arXiv: 2602.09784v1 [cs.LG], 10 Feb 2026
Overview
- Research area: Mechanistic interpretability of transformer language models, specifically the intersection of circuit discovery and activation steering.
- Technical level: Advanced. The paper assumes familiarity with attention heads, MLPs, the residual stream, activation patching, attribution patching, and the linear representation hypothesis.
- Scope: The paper argues and tests the claim that a single geometric object — directions encoded by answer tokens processed in isolation — can both identify which components belong to a task circuit and steer model behavior toward a target output.
What This Paper Is About
Circuit discovery (finding which attention heads and MLPs cause a behavior) and activation steering (controlling behavior by adding directions to activations) are usually treated as two separate research threads, even though they operate on the same internal representations. The authors propose the Circuit Fingerprint hypothesis: answer tokens, when passed through the model in isolation, trace the same computational pathways that produce them, so their representation differences carry a geometric signature of the circuit. The goal is to show that reading this signature recovers circuit structure without gradients or causal intervention, and that writing along the same directions produces controlled behavior changes.
Key Contributions
- Circuit membership can be read geometrically. The authors show that alignment between contrastive prompt activation differences and answer-token-derived directions recovers circuit structure comparable to gradient-based methods (EAP, EAP-IG-inputs) without requiring backpropagation.
- Read and write are dual operations. The same directions used to identify circuit components also enable steering, which the authors present as causal validation that they are manipulating genuine computational pathways rather than superficial correlations.
- Feature circuits can be discovered from instructions alone. By extracting directions from instruction-modified prompts (for example, "Answer with joy: {prompt}"), the method identifies feature-relevant heads without any task-specific dataset.
- A Shapley-based channel decomposition. The method assigns Query, Key, and Value contributions within each attention head using Shapley values over a three-player cooperative game, which the authors report clusters heads by functional role.
Main Findings
- Circuit discovery without gradients works comparably. Evaluated on IOI, SVA, and MCQA across GPT-2 Small, Qwen2.5-0.5B, Llama 3.2-1B, and OPT-1.3B, the proposed Circuit Fingerprint (CF) method reports CMD and CPR values in the same range as EAP and EAP-IG-inputs. For example, on GPT-2 Small IOI, CF reports CMD 0.06 and CPR 0.98, versus EAP at CMD 0.03 / CPR 0.97 and EAP-IG-inputs at CMD 0.03 / CPR 0.97. The authors state their method is often very similar to EAP but lags behind EAP-IG-inputs, which they attribute partly to simplifications in their implementation. MCQA is not used for GPT-2 Small because, in the authors' words, there is no circuit in the model for it.
- Larger models score better on CMD and CPR. The authors suggest this is related to better disentanglement of concepts in larger models.
- The method locates the same critical heads as gradient methods. On the IOI prompt "After Mary and Bob went to the store. Bob gave a bottle to", both the geometric identity scores and EAP-IG-inputs identify the same critical heads in layers 9–11, with the strongest signal at the final token position.
- An incidental observation about token "gave". The simple analysis shows the token "gave" exhibits a similar activation structure to the final token "to", which the authors interpret as the model having had "Mary" as an option to output after it.
- Shapley decomposition clusters heads by function. In GPT-2 Small, 15 attention heads are known to be critical to the IOI circuit. Splitting heads by the sign of their total contribution, negative heads include Duplicate Token Heads, Negative Name Movers, and Previous Token Heads, while positive heads include S-Inhibition Heads, Name Movers, and Induction Heads. In a ternary plot of Q/K/V importance at the final token, output-focused heads (Name Movers) are Q-dominated while routing heads (S-Inhibition) are K-dominated.
- Steering matches activation patching closely on IOI. On 100 examples, at full intervention strength (α = 1), answer token directions yield P(correct) = 0.014 versus 0.0 for activation patching, with logit differences of −4.07 versus −7.34. The authors note the slightly weaker effect suggests their directions capture the discriminative signal without fully replicating the corrupted distribution.
- Emotion steering beats instruction prompting. Geometric steering reaches 69.8% emotion classification accuracy versus 53.1% for instruction prompting, with median perplexity 13.37 versus 17.03, and factual accuracy 89.6% versus 90.1%.
- Per-emotion results are uneven. Geometric steering achieves higher classification accuracy on 3 of 5 emotions (joy, sadness, surprise). Joy: 87.8% steered versus 49.0% baseline; sadness: 89.8% versus 53.1%; surprise: 79.6% versus 32.7%. It underperforms on disgust (67.3% versus 75.5%) and anger (24.5% versus 55.1%). Factual accuracy under steering was joy 100%, anger 96%, surprise 93%, sadness 81%, disgust 78%.
- Factual degradation is valence-dependent. Positive-valence steering (joy) preserves or improves factuality at 100%, while negative-valence emotions degrade it (sadness 81%, disgust 78%), producing errors such as "Alexander Bonniweeper" for Bell and "Albert Sissoar" for Einstein. The authors suggest negative emotion representations may be less disentangled from lexical content in activation space.
Methodology in Plain English
The method uses two information sources. First, answer tokens: the correct answer ("Paris") and an incorrect one ("Rome") are each passed through the model in isolation, and their representation difference at the last layer defines a target direction. This takes two forward passes. Second, contrastive prompts: a clean prompt ("The capital of France is") and a corrupted one ("The capital of Italy is") are run, and the activation difference between them shows how the model's internal state changes.
Circuit components are then identified by how much their contrastive-prompt activation difference aligns with the answer-derived target direction. Crucially, alignment is measured in each component's native space rather than the shared residual stream, because projecting through the output matrices (W_O for attention, W_out for MLPs) mixes the component's computation with shared geometry. The authors show this decomposition preserves additivity: component scores sum to the total projection onto the target direction.
For edges, the method decomposes a downstream component's task-relevant input into contributions from each upstream component, giving per-channel ratios for Query, Key, and Value that sum to one. To weight Q, K, and V when aggregating a head's edge importance, the authors treat the three channels as players in a cooperative game and use Shapley values, which requires evaluating 2³ = 8 coalitions per head. Total component importance is then a component's direct importance plus its indirect importance through downstream edges, computed in a backward pass through the model. The authors state their goal is not an optimal discovery method but a demonstration that geometric alignment recovers circuit structure; they simplify by focusing on the final token position and ignoring indirect effects from earlier positions and layer normalization.
For steering, the authors build an intervention subspace from answer-token representations related to a source and target feature, center them, and compute an orthonormal basis via SVD. When the target token is known (factual recall), the intervention uses the separation between source and target as the scale; for stylistic features (emotion, language) where the target token is unknown, the source magnitude is transferred instead. Interventions are applied at the attention heads identified by circuit discovery. In the steering experiments they use 100 English prompts from Konen et al. (2024), split into factual and subjective categories, and steer Llama3.2-1B toward five emotions: joy, anger, sadness, surprise, and disgust. The number of attention heads (25) is the only hyperparameter, chosen based on performance. Emotion classification accuracy is measured with a DistilRoBERTa model fine-tuned on emotion detection, and perplexity is measured against GPT-2 Large.
Why This Matters
- Impact on research: The paper connects circuit discovery to the linear representation hypothesis, arguing that if features are directions in activation space, then circuits manipulating those features should be identifiable through directional alignment. If correct, this simplifies the interpretability toolkit: instead of separate methods for discovery and control, a single geometric analysis reveals both which components matter and how to manipulate them. The authors frame circuits as fundamentally geometric structures rather than only causal ones.
- Real-world applications (as suggested by the method's scope):
- Controllable text generation, steering outputs toward a target emotion or style directly rather than through prompting.
- Persona and style adaptation, where instruction prefixes such as "Answer as a pirate" identify the heads governing that style.
- Language steering, mentioned by the authors as an extension of the same template used for emotions.
- Factual recall editing, where knowledge directions are known, using the source-to-target separation as the intervention scale.
- Industry relevance: Applications that need precise behavioral control of deployed models — tone moderation, persona consistency, output style enforcement — could potentially use direct circuit intervention instead of prompt engineering. The paper reports that geometric steering outperforms instruction prompting on emotion classification (69.8% versus 53.1%) while keeping factual accuracy nearly unchanged (89.6% versus 90.1%), though it also documents cases where steering degrades factual recall.
Future Directions
- Position-level effects beyond the final token. The current implementation focuses on the final token position; extending to earlier positions is listed as future work.
- LayerNorm nonlinearities in edge attribution. The authors explicitly note their edge attribution ignores layer normalization and call for addressing it.
- Preserving factual grounding during steering. The observed factual degradation for negative-valence emotions (81% for sadness, 78% for disgust) and for distant language pairs is flagged as an open problem.
- More robust persona-based circuit discovery. Zero-shot steering remains fragile, and the authors state that even grammatically correct outputs can exhibit semantic degradation, suggesting further work on disentangling features from semantic content.
Target Audience
This paper is most useful for mechanistic interpretability researchers already familiar with circuit discovery and activation steering, and for machine learning engineers working on controllable generation who want an alternative to gradient-based attribution and prompt-based control. Readers without background in transformer internals, the residual stream, and attribution methods will find the terminology dense; the paper is written at an advanced technical level and assumes prior exposure to work such as activation patching, EAP/EAP-IG, and the linear representation hypothesis.
Authors’ abstract
Circuit discovery and activation steering in transformers have developed as separate research threads, yet both operate on the same representational space. Are they two views of the same underlying structure? We show they follow a single geometric principle: answer tokens, processed in isolation, encode the directions that would produce them. This Circuit Fingerprint hypothesis enables circuit discovery without gradients or causal intervention -- recovering comparable structure to gradient-based methods through geometric alignment alone. We validate this on standard benchmarks (IOI, SVA, MCQA) across four model families, achieving circuit discovery performance comparable to gradient-based methods. The same directions that identify circuit components also enable controlled steering -- achieving 69.8\% emotion classification accuracy versus 53.1\% for instruction prompting while preserving factual accuracy. Beyond method development, this read-write duality reveals that transformer circuits are fundamentally geometric structures: interpretability and controllability are two facets of the same object.