Research
The Geometry of Inference in Transformer Residual Streams
Overview Research area: Mechanistic interpretability and representation geometry in transformer language models — specifically how the residual stream progressively becomes specific to a particular ne

- arXiv
- 2609.37824
- Published
- 2026-09-29
- Authors
- Timur Mudarisov, Mikhail Burtsev, Radu State
AI summary
Overview
Research area: Mechanistic interpretability and representation geometry in transformer language models — specifically how the residual stream progressively becomes specific to a particular next-token prediction.
Technical level: Advanced. The paper combines large-scale empirical probing of six pretrained models with formal derivations (margin algebra, a nested-competitor-set proof, a high-dimensional shared-axis model, and spherical/von Mises–Fisher distribution fitting).
Scope: A single study measuring how intermediate residual states in six pretrained language models distinguish their own final state from a bank of final states drawn from other contexts, plus theoretical models explaining when and why that distinction sharpens with depth.
What This Paper Is About
Intermediate layers of a transformer contain information about the token the model will eventually predict, but the paper asks a narrower question: how selectively does an intermediate state distinguish the final representation it will actually reach from the final representations reached by other contexts? To measure this, the authors treat each context's own final residual state as its "own endpoint" and all other contexts' final states as a reference bank of "alternatives," counting how many alternatives sit closer to the current state than the own endpoint does. The goal is to characterize how this geometric specificity grows across depth, why it can improve even when the Euclidean distance to the own endpoint stays flat, and how the shape of the endpoint population governs the process.
Key Contributions
-
Empirical characterization across six models. The authors relate cosine endpoint distance to output-token rank, showing that endpoints whose top predictions are less probable under the query context lie farther away, and that an early average preference for the own endpoint coexists with many individual closer alternatives. Cosine competitor sets show entries as well as exits over depths where mean competitor count declines, while Euclidean sets are closer to nested and change mainly near the end.
-
A theoretical account of competitor counts. The paper relates expected competitor counts to endpoint probability mass (the "competitive mass" of a comparison ball or directional cap), shows how gradual directional alignment can sharply reduce competition during a plateau in Euclidean distance, and proves that monotone straight-line convergence to the own endpoint yields nested competitor sets under both Euclidean and cosine distance.
-
A test of whether endpoint geometry alone predicts competition. Structured endpoint distributions fitted only to final states predict mean competitor fractions along held-out residual trajectories more accurately than a uniform-sphere baseline, with the projected-normal model lowest in five of six models.
-
Controls isolating the source of early preference. Excluding endpoints from the query's own document leaves the cosine preference and competitor-fraction curves essentially unchanged; restricting alternatives to contexts ending in the same input token preserves preference from the first measured layer with smaller margins and more competitors; removing each state's component along its input-token embedding direction preserves first-layer cosine preference in all six models.
Main Findings
-
Average preference appears early. Across all six models and both distance metrics, the document-averaged margin is positive from the earliest measured post-block state (margin defined as the mean alternative distance minus the own-endpoint distance, so positive means the own endpoint is closer than the average alternative).
-
Early preference is not selective. Many individual competitors remain after that preference becomes detectable. Under Euclidean distance, five models retain hundreds of competitors across substantial portions of their depth.
-
Lower-ranked output tokens correspond to farther endpoints. Across all six models, endpoints whose top predictions receive lower probability under the query context tend to lie farther from the query's own endpoint in cosine distance, with positive trends in one-sided tests. Endpoints sharing a top prediction are also more compact on average, and greater endpoint distance is associated with greater divergence between output distributions.
-
Competitor sets shrink but change membership. Adjacent-layer Jaccard overlap plus entry and exit fractions show both entries and exits under cosine distance, including sustained turnover in Qwen2.5-1.5B. Entries occur over depths where the mean competitor count decreases, which excludes a universally nested elimination process. Euclidean sets are closer to nested over much of the trajectory.
-
Directional progress can accompany flat distance. Cosine competition often decreases earlier than Euclidean competition, and directional alignment can improve while Euclidean endpoint distance remains nearly constant, particularly in intermediate layers of Mistral-7B and Llama-3-8B. Cosine profiles persist under uniform subsampling of the endpoint bank.
-
Fewer competitors does not mean tighter competitors. The paper separates set size (comparisons with the residual state) from cohesion (distances among the endpoints themselves) and gives a fixed-bank construction in which the count falls while mean pairwise Euclidean and cosine distances both increase. The authors state they do not infer an empirical cohesion trend from the turnover or whole-layer spread measurements.
-
Straight-line convergence is ruled out where entries occur. For a straight segment from an initial state to the own endpoint, both margins are affine and nonnegative at the endpoint, so any alternative outside the competitor set stays outside. Observed endpoint entries therefore establish departures from monotone straight-line convergence on the affected trajectories.
-
A high-dimensional model explains sharp count drops. Decomposing endpoints and states into a shared direction plus an orthogonal subspace, an alternative competes exactly when its orthogonal alignment exceeds the own endpoint's. Any positive alignment makes the own endpoint closer than a random alternative on average, but when alignment is small the expected competitor fraction stays close to one half; an expected count much smaller than one requires the competitive mass to fall below 1/M. The rate of count reduction is governed by the ratio of the alternative-score density to the competitive mass, times the alignment speed.
-
Norm changes allow plateaus. Setting the norm ratio to twice the alignment keeps normalized Euclidean distance equal to one while alignment increases, which makes a flat Euclidean curve compatible with directional refinement.
-
Fitted distributions predict competition. The projected-normal family achieves the lowest prediction error on cosine competitor fractions in five of six models; the von Mises–Fisher mixture is slightly better for Gemma-7B and ranks second elsewhere. Both outperform the spherical cap and single von Mises–Fisher models across all six architectures, and the uniform sphere has the largest error in five. Best held-out diagnostic scores from endpoint-distribution fitting improve over the sphere by factors of 4.6 to 7.5.
Methodology in Plain English
The researchers take six pretrained language models — Gemma-2B and Gemma-7B, Qwen2.5-1.5B and Qwen2.5-7B, Mistral-7B, and Llama-3-8B — and run each on 1,024 non-overlapping 256-token contexts drawn from FineWeb sample-10BT. For every context they record the residual state at the final input position after each block, along with the final state before final normalization. That final state is the context's "own endpoint."
For each context, the reference bank consists of all other endpoints, giving 1,023 alternatives. At each depth they compare the current state against the own endpoint and against every alternative using two distances: Euclidean distance divided by the own endpoint's norm, and cosine distance. Any alternative closer than the own endpoint is counted as a competitor; the competitor fraction is the count divided by the bank size, and the own endpoint's strict retrieval rank is one plus the competitor count. To connect geometry to prediction, they rank the model's output tokens by final probability and measure the mean cosine distance from the query's endpoint to endpoints whose top prediction is that ranked token.
To study whether refinement removes competitors or reshuffles them, they compute adjacent-layer Jaccard overlap along with entry and exit fractions, normalizing entry fractions by the current set size and exit fractions by the previous set size before averaging across contexts.
Theory work runs alongside the measurements. Signed margins express each comparison, and their affine dependence on the state yields both the update identities and the straight-path nesting proof. A shared-axis model with a fixed shared direction, a tunable alignment parameter for endpoints, and a tunable alignment for intermediate states gives closed-form expressions for expected competitor counts, the competitive mass of the comparison region, and the rate at which counts decline.
Finally, they fit five families of distributions to the final endpoints alone — uniform sphere, spherical cap, single von Mises–Fisher, von Mises–Fisher mixture, and projected normal — splitting documents roughly 60/20/20 into training, validation, and test sets and refitting on the combined training and validation endpoints. All families use the same empirical distribution of endpoint norms sampled independently of direction, so differences reflect directional structure. They then sample banks of 256 endpoints from each fitted distribution and evaluate predicted mean competitor fractions along held-out trajectories while keeping intermediate states and own endpoints fixed, averaging over eight banks. Because competitor fractions span several orders of magnitude, comparison uses mean absolute difference on a base-10 logarithmic scale with an offset of 10 to the minus 3, and the final layer is excluded because competitor fractions are zero there by construction.
Why This Matters
The paper reframes a familiar interpretability observation — that intermediate layers already encode the eventual prediction — as a question about relative geometry among final states. It supplies both a measurement protocol (a fixed endpoint bank with competitor counts) and a theory that separates three things often conflated: Euclidean distance to the target, directional alignment with it, and the number of nearby alternatives. It also offers a concrete negative result: because straight-line convergence provably yields nested competitor sets, observed competitor entries are direct evidence that residual trajectories are not simply homing in on their endpoint.
The authors are explicit about limits. Endpoint geometry explains only part of the variation in model outputs, the competitor fraction is not a calibrated token probability, competitor sets depend on the chosen bank and metric, and their size reaches zero at the own endpoint by construction. The observed turnover does not establish an explicit internal search process inside the model.
Potential applications, none of which are demonstrated or tested in this paper:
- Interpretability tooling, where competitor counts and turnover statistics could serve as depth-resolved diagnostics of how a specific model narrows toward a prediction.
- Model comparison and auditing, since the depth profiles and turnover behavior differ across architectures and could serve as structural fingerprints.
- Representation monitoring, where margin crossings between an intermediate state and nearby alternative endpoints could mark the depths at which predictions are most changeable.
- Compression or early-exit research, since the theory identifies when alignment improves without distance improvement, which is relevant to deciding which layers carry directional refinement versus scale change.
Industry relevance: The work targets practitioners who probe, monitor, or compare transformer internals rather than those training models from scratch. The measurement framework is model-agnostic and requires only stored residual states and a bank of final states, but it is an analysis method, not a system that improves downstream accuracy.
Future Directions
-
Tracking development across training. The fitted distributions predict competition along observed trajectories but leave the origin and timing of residual updates unexplained. The authors propose following endpoint populations and residual trajectories across training checkpoints to see how the two organize together.
-
Intervening at margin crossings. The paper suggests interventions near competitor-margin crossings to test whether changes in geometric preference actually affect token predictions, since the current results are correlational.
-
Testing whether turnover reflects search. The observed turnover does not establish an explicit internal search process, and the authors leave the functional role of geometric competition open.
-
Extending the endpoint model to update dynamics. The endpoint distribution is fitted only to final states; the paper does not model which residual updates produce the observed alignment schedules, and it notes that the depth at which counts fall depends on the prescribed alignment trajectory.
Target Audience
This paper is aimed at mechanistic interpretability and representation-geometry researchers who are comfortable with high-dimensional probability, margin algebra, and distribution fitting, and who already work with residual-stream probes. It will also interest interpretability-adjacent engineers building tooling that inspects intermediate states across depth, and theorists studying how transformer representations converge toward their outputs. Readers looking for benchmark accuracy improvements, training recipes, or applied deployment guidance will not find them here, and no single aggregate performance metric is reported for the six models.
Authors’ abstract
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.