Research
The information geometry of large language models is shared, learned, and controllable
Overview Research area: Machine learning theory and interpretability — specifically, the information geometry of next-token probability distributions in large language models, with connections to mech

- arXiv
- 2609.11063
- Published
- 2026-09-10
- Authors
- Dario Picozzi
AI summary
Overview
Research area: Machine learning theory and interpretability — specifically, the information geometry of next-token probability distributions in large language models, with connections to mechanistic interpretability, model steering, knowledge editing and fine-tuning.
Technical level: Advanced. The paper assumes familiarity with Riemannian geometry, Fisher information, natural-gradient methods, sparse autoencoders and transformer internals, though the results are stated as measurable, testable relationships rather than purely formal proofs.
Scope (one sentence): The paper argues that the Fisher–Rao geometry of a model's output distribution is a canonical, behaviour-determined structure that is shared across architectures, learned from corpus statistics, and directly usable to prescribe minimum-disturbance interventions.
What This Paper Is About
A language model's internal activations can be re-coordinated arbitrarily — the same input–output function can be realised under any invertible linear change of hidden-state coordinates — so Euclidean measures of activation distance, feature importance and steering strength are not invariant. The paper's central claim is that the model's next-token probability distribution, unlike its activations, does have a privileged geometry: the Fisher–Rao metric, which Chentsov's theorem identifies as the unique canonical metric on distributions up to overall scale.
The goal is to show that this output geometry is (i) identifiable from predictive behaviour alone, (ii) shared across independently trained models of different architectures, (iii) shaped by and predictable from corpus statistics, and (iv) the right object to use when changing one behaviour while leaving others intact.
Key Contributions
-
A behavioural identifiability result plus an empirical test. Under stated finite-support and inverse-chart conditions, held-out profiled predictive risk obeys a global quadratic margin
K(S) ≥ c·d²(S,T), so zero profiled risk identifies the language-defined read-out subspace exactly at matched rank. This is tested on Pythia 70M–1.4B, where profiled loss grows with squared subspace distance with R² ≥ 0.9999 on all eight paths and steeper slopes for real read-out deviations than for matched-random controls (slope ratios 1.14–1.27). -
Cross-model and cross-architecture sharing of output geometry. Across ten models spanning transformer, state-space and recurrent architectures, seven training pipelines, four tokenizers and 70 million to 7 billion parameters, output Fisher–Rao geometries agree at a mean rank agreement of 0.88, against 0.62 for mid-layer and 0.61 for last-layer activation geometries, with a bootstrap interval excluding zero. A byte-level, tokenizer-independent check gives cross-tokenizer agreement of 0.91.
-
Language statistics predict the geometry and when behaviours are acquired. Token probabilities and learned read-out structure jointly predict the output-Fisher spectrum and its effective dimension without fitted parameters, and ex-ante corpus n-gram margins predict acquisition trajectories of fresh facts without refitting (trajectory R² = 0.775–0.792).
-
A minimum-disturbance intervention framework. The damped natural-gradient solve is shown to be the unique minimum-regularised-cost step at a fixed objective change, with a no-fitted-parameter prediction of its advantage over Euclidean control; the same correction is reported to improve steering, knowledge editing, attribution, dictionary learning and fine-tuning.
Main Findings
-
Behaviour fixes the output geometry up to scale. The pullback Fisher metric
G_h = J_hᵀ H_h J_hpredicts realised output change viaKL ≈ ½ δhᵀ G δh. This matched measurement across 99 model–depth–objective cells spanning 11 models, six families and 125M–1.5B parameters, with a median measured-to-predicted KL ratio of 1.000 (range 0.948–1.224). -
The natural-gradient direction is coordinate-covariant. Transporting it across a seeded family of activation reparameterisations reproduces the native direction to below 10⁻¹², whereas resetting the reference metric
Rto the identity changes it by order one. -
Activation geometry retains an exact GL(d) gauge freedom. Invertible reparameterisations of a hidden layer, compensated downstream, leave every conditional law unchanged while altering Euclidean activation geometry, so the resolved output geometry is invariant under behaviour-preserving reparameterisation.
-
The assigned language law, not the architecture, sets the geometry. Across eight pairs of synthetic languages with identical token frequencies and conditional entropy, changing the assigned law recovers more than 99% of the imposed squared geometric separation at matched predictive accuracy, while changing architecture within a law leaves geometry nearly unchanged.
-
Cross-tokenizer agreement holds in a common outcome space. Mapping each next-token law to the distribution of the first byte of the remaining text gives cross-tokenizer agreement of 0.91, and coarse-graining contracts the divergence on every pair. Partialling out orthographic surface geometry or the strongest corpus n-gram predictor leaves at least 98% of the agreement in place.
-
Agreement is quantitatively accountable. A six-term variance/covariance decomposition reconstructs every measured configuration to numerical precision across 112 model-pair configurations. Fitted on half the contexts, its exchangeable restriction predicts held-out agreement with pooled median absolute error 0.009 across 96 defined cells; from step 256, checkpoint-median errors are below 0.014.
-
What is shared carries semantic alignment. Consensus geometry aligns with an independent sentence-encoder semantic geometry on every battery, rising to 0.44 when the next token answers a factual relation and surviving a surface control at +0.23; residual alignment is indistinguishable from zero on the natural and templated batteries and small (+0.05) on the semantic battery. A linear semantic probe trained on one model transfers across models with eight-way accuracy 0.66 (0.70 through the consensus geometry alone), versus 0.72 within-model and 0.125 chance, while transfer through the residual is below chance. Residual eigenvector sharing between families is about 9% of the raw overlap against a displacement-matched null.
-
Model geometry approaches human completion structure with scale and fit. On 512 sentence contexts with human completion norms, mean squared distance falls from about 0.34 at 70M to 0.30 at 2.8B parameters, with reliability-corrected relational alignment of about 0.82–0.85 across six sizes. On a separate 384-context battery, a risk-to-alignment relation fitted on Pythia predicts OLMo checkpoints and five external models without refitting (correlation 0.758, RMSE 0.0295).
-
Directional alignment with human structure and calibration gains. On a corpus of 1,726 human next-word predictions mapped to the same 65 outcomes, model-derived leading directions recover human directional structure above a random subspace control, with positive simultaneous lower bounds in six of seven model families. A calibration learned only from model laws improved expected human log score in six of seven families, with gains 0.0196–0.0507 nats (Qwen2: −0.0030 nats).
-
A scale trend replicated on independent data. On 640 word positions in five previously unused narratives (64,000 responses), semantic distance fell with log parameter count (slope −0.00872, 95% interval −0.01202 to −0.00556) and fell from initial to final checkpoint in both Pythia and OLMo.
-
Spectra follow an approximately inverse-rank law. Output-Fisher eigenvalues behave as λᵢ ∼ i⁻ᵝ with β ≈ 1 over the rank-2–80 core window from 70M to 6.9B parameters across families and vocabularies; rank-8–32 exponents remain of order one across all nine models from 125M to 1.5B parameters in the external-family panel, tracking the model's own weighted token profile (median deviation 0.04, r = 0.91).
-
Effective dimension is predictable without fitted parameters. The weighted profile
N̂_eff(α) = Σ q₍ₖ₎/(q₍ₖ₎ + α)predicts held-out spectra for a disjoint 64-context battery in nine models from five external families (BLOOM, GPT-Neo, Mamba, Qwen2, RWKV): rank-8–32 family-median eigenvalue errors of 0.0617–0.0884 decades, multiplicative deviations of 1.15–1.23-fold, spectral-exponent errors of 0.0623–0.0856, and N_eff curve RMSEs of 0.0270–0.0521. Errors are 4.89–9.22-fold lower than a same-trace flat spectrum and 13.55–25.63-fold lower than a rank-shuffled profile; effective-dimension error is 6.44–7.88-fold lower than the flat-spectrum curve. Exact inheritance certificates apply in 98.4–100% of external-family cells. -
Probability concentration and read-out structure play distinct roles. In a matched 14-model, seven-family comparison on 192 fresh contexts, a probability-only profile predicts effective dimension more accurately in every family, while the weighted profile predicts the spectral exponent more accurately in every family, with exact one-sided family sign-flip P = 1/128 for each contrast.
-
Ex-ante corpus statistics predict acquisition. N-gram margins measured before training at unigram, bigram and trigram levels predict, without refitting, complete trajectories of fresh facts disjoint from calibration: trajectory R² = 0.775–0.792, acquisition-status agreement 0.792–0.854, and continuous timing error of 0.77–0.96 log₂ training-step units. Figure 4a reports per-size values of 0.790, 0.792 and 0.775 for trajectory R²; 0.792, 0.854 and 0.807 for status agreement; and 0.77, 0.95 and 0.96 log₂ units for timing error at Pythia 70M, 160M and 410M.
-
Evidence depth causally delays acquisition. In a randomised synthetic-language test with paired initial weights and batch streams, acquisition-time quantiles under deep evidence were approximately 4.3 times the corresponding shallow-evidence quantiles across the analysed 10th–20th percentiles, a near-uniform shift of 2.10 in log₂ training steps, stable under leave-one-seed-out analysis. In a fully crossed randomised study of GPT-2-, GPT-NeoX- and Llama-style decoders and two controlled languages, deeper evidence delayed persistent acquisition in every cell by 0.81–1.51 log₂ training-step units (1.75–2.86-fold more training steps), with all six 95% intervals excluding zero and all 12 control intervals containing zero.
-
Acquisition curves align better by loss than by step. Aligning cumulative acquisition curves for the same 300 facts by held-out loss gives a median absolute error of 2.1 facts, about three times lower than alignment by training step, in the descriptive 70M-to-160M comparison.
-
Intervention cost is predicted without fitted parameters. The finite-damping cost ratio
R_pred(q) = Π_G(q)/Π_G(A⁻¹q), fixed before intervention, tracks 3,515 usable matched-response measurements from 1,188 cells (98.6% coverage) with median measured-to-predicted ratio 0.967 (95% interval 0.957–0.975) and log–log slope 0.982 (95% interval 0.966–0.997) across 11 models, six families, three objectives and three relative layer depths. -
Fisher paths accumulate far less off-target divergence. With both methods relinearised after every accepted step, 864 points from 216 prompts, 16 models and 18 model–objective cells give 72 cell–target geometric-mean cost ratios spanning 11.5–127.1, with every paired 95% interval above one.
-
Updates are reusable across prompts. A shared update learned from four donor prompts per objective, using the Fisher metric averaged over four reference prompts, transfers a positive mean change in all 24 model–objective–target settings on BLOOM-560M, Pythia-410M, Pythia-1.4B and Qwen2-1.5B, without computing target gradients. At matched donor changes it reduces reference-continuation sequence KL by roughly 3–6 times relative to Euclidean control and at least 2 times relative to a donor-only Fisher metric (mean Euclidean/reference-Fisher ratios 2.90–6.29).
-
Prompt-specific control extends to instruction-tuned models. At equal aggregate mean objective change, natural-gradient frontiers have 8–74 times lower off-target divergence for anti-sycophancy, 30–190 times for capital-city truth preferences and 4–550 times for style.
-
One geometric correction spans operations. On 30 held-out CounterFact records with GPT-2, the pullback-metric edit had 10.57-fold lower geometric-mean other-token divergence than its Euclidean counterpart (95% interval 8.52–13.37), with both operators attaining the fixed +5 log-odds shift on all 30 edits; separate GPT-2-XL evaluations cover 100 edits under standard metrics and 300 under a stricter probability-margin criterion. Across 303 zero ablations of sparse-autoencoder features at Pythia-410M, Fisher cost rank-correlated with exact divergence at ρ = 0.997 versus 0.896 for activation magnitude. A code-Gram preconditioner raised Hungarian-matched mean absolute decoder–ground-truth cosine from 0.167 to 0.480 (+0.313). Natural-gradient low-rank adaptation produced 60–250-fold less off-target change than Adam at matched on-target change at 410M, and improved the preservation-KL frontier 1.57-fold relative to tuned AdamW with explicit KL regularisation at 70M.
-
The available correction shrinks with depth and magnitude. The advantage is largest at early-to-mid injection layers and shrinks toward the near-linear read-out; it vanishes in the isotropic limit, where natural-gradient and Euclidean steps coincide. Model-balanced geometric means over the three declared intervention fractions are 0.918, 0.860 and 0.755.
Methodology in Plain English
The approach treats a model's output — its probability distribution over the next token — as the object of study, rather than its internal activations. Distances between output distributions are measured with the Fisher–Rao metric, and that metric is "pulled back" through the network to score how much a small change to an internal activation would change the model's behaviour. Because output distributions are invariant to any coordinate change in the hidden states, this gives a comparison scale that does not depend on architecture, tokenizer or activation alignment.
To test whether models really share this structure, the authors have each model produce a matrix of pairwise Fisher–Rao distances between its own next-token distributions over a common battery of contexts, then compare matrices using Spearman rank correlation between their upper triangles. This needs no shared vocabulary and no activation alignment. The same comparison is repeated after mapping laws to a byte-level outcome space to rule out tokenizer artefacts.
For the "learned" claim, they train transformer, gated-recurrent and diagonal-recurrent models at two capacities on each of eight pairs of synthetic languages with known conditional laws, identical token frequencies and conditional entropy — a controlled experiment in which the assigned law is the only thing that changes. For acquisition timing, they use randomised synthetic languages where otherwise identical facts are assigned to shallow or deep evidence constructions with paired initial weights and batch streams, plus a fully crossed randomised study across three decoder styles and two controlled languages.
For interventions, they solve for the minimum-cost step under the damped natural-gradient rule, derive a no-fitted-parameter prediction of its advantage over a Euclidean step, and then measure outcomes at matched behaviour change. Tests span 11 models and six families for steering, 30 CounterFact edits for knowledge editing, 303 sparse-autoencoder feature ablations at Pythia-410M for attribution, and sparse-autoencoder decoder training plus low-rank adaptation for the training-time claims.
Why This Matters
Impact on research. The paper reframes the "Platonic representation hypothesis" debate: instead
Authors’ abstract
Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.