Skip to content
AI.info

Research

The Information Geometry of Softmax: Probing and Steering

Overview Research area: interpretability and representation geometry for softmax-based AI models (language and vision-language). Technical level: Advanced (relies on information geometry, Bregman dive

The Information Geometry of Softmax: Probing and Steering
arXiv
2602.15293
Published
2026-02-17
Authors
Kiho Park, Todd Nief, Yo Joong Choe, Victor Veitch

AI summary

Overview

Research area: interpretability and representation geometry for softmax-based AI models (language and vision-language). Technical level: Advanced (relies on information geometry, Bregman divergences, and convex duality, though the intuition is explained in plain terms). Scope: the paper argues that the natural geometry of softmax representation spaces is information geometry, and uses that geometry to derive "dual steering," a method for manipulating concepts via linear probes while preserving off-target behavior. Authors are Kiho Park, Todd Nief, Yo Joong Choe, and Victor Veitch (University of Chicago; INSEAD). Code is available at github.com/KihoPark/dual-steering.

What This Paper Is About

The linear representation hypothesis says high-level concepts correspond to directions in a model's representation space, but methods built on it are often brittle, and the paper traces part of the problem to an implicit assumption that representation space has a flat, Euclidean geometry. For representations that define softmax distributions, the authors argue the correct notion of "closeness" is closeness of the induced probability distributions, which is exactly information geometry. They use this framework to study interpolation between representations and to derive a new steering method, dual steering, that modifies a target concept while minimizing unintended changes to off-target concepts.

Key Contributions

  1. Identifies the natural geometry of softmax representation vectors as a Bregman (dually flat) geometry, induced by the log-normalizer A, and shows the resulting duality structure (primal Λ and dual Φ spaces, linked by φ(λ) = ∇A(λ)) is central to how semantics are encoded.
  2. Shows there are two distinct natural interpolations between representation vectors — the primal e-geodesic and the dual m-geodesic — which produce distinct semantic behavior (primal acts like an AND/intersection operator; dual acts like an OR/union operator), and argues a flat geometry cannot capture this structure.
  3. Introduces dual steering, which adds the probe vector in the dual space rather than the primal space, and proves (Theorem 3.1) that it optimally modifies the target concept while minimizing changes to off-target concepts under a concept-factorizable distribution assumption.
  4. Empirically validates dual steering on Gemma-3-4B and MetaCLIP-2, showing improved controllability and stability relative to standard Euclidean steering across three robustness metrics.

Main Findings

  • Softmax geometry is Bregman geometry. The KL divergence between softmax distributions equals a Bregman divergence induced by the log-normalizer A, giving a dually flat structure with a dual map φ(λ) = ∇A(λ) = E[γ | λ] and an inverse defined via the convex conjugate A*.
  • Primal and dual interpolation behave differently. The primal (e-geodesic) interpolation λ_t = (1−t)λ_0 + tλ_1 minimizes a weighted sum of reverse KL divergences and emphasizes shared structure (e.g., for "a black dog" vs. "a white dog," the probability of a black-and-white dog image rises at the midpoint). The dual (m-geodesic) interpolation φ_t = (1−t)φ(λ_0) + tφ(λ_1) minimizes a weighted sum of forward KL divergences and corresponds to a linear mixture of the endpoint distributions.
  • A type mismatch underlies Euclidean steering. The probe vector β_W is an element of the dual space (a linear operator on Λ), but standard Euclidean steering adds it directly to a primal representation λ_0, which is only valid when primal and dual spaces coincide.
  • Dual steering preserves off-target distributions. Theorem 3.1 shows that under a linear probe P(W=1|λ) = σ(β_W^T λ + b_W), the minimizer of D_KL(P_{λ_0} ‖ P_λ) over the hyperplane Λ_W(c) satisfies φ(λ̂) = φ(λ_0) + t β_W, i.e., it is a dual steering step. When distributions are concept-factorizable, this minimizer also minimizes D_KL(P^Z_{λ_0} ‖ P^Z_λ) — preserving the off-target distribution.
  • Euclidean steering is asymmetric. For a probe in the dual coordinate, the reverse KL projection decomposes into an on-target term plus an off-target term whose coefficient, the total mass on counterfactual pairs, depends on λ and cannot be ignored. This produces the "leakage" seen in practice, where mass shifts to unrelated tokens such as "friend" or "to."
  • Practical feasibility requires care. The dual space Φ is only a bounded convex set (the convex hull of the unembedding vectors), and the update system Cov[γ|λ] Δλ = ε β_W becomes rank-deficient when the softmax is highly concentrated. A regularized Newton update, (Cov[γ|λ] + α I_d) v = ε β_W, restores invertibility, and solving via conjugate gradient gives complexity O(nkd) instead of the naive O(kd² + d³), with n = 20 conjugate gradient iterations.
  • Dual steering wins on all three robustness metrics. Across Gemma-3-4B and MetaCLIP-2 tasks, both methods raise the target concept probability, but Euclidean steering distorts the off-target distribution while dual steering maintains it — measured by total mass on counterfactual pairs (constant is better), KL divergence of off-target distributions, and weighted inverse rank differences (lower is better).
  • When Euclidean steering happens to work. It tends to preserve off-target distributions when the counterfactual sum stays roughly constant, which often occurs with the Primal MD direction; but even then the ranking of off-target components can shift substantially, so dual steering remains more robust.

Methodology in Plain English

The authors start from the observation that any representation vector λ defines a probability distribution over a set of items via softmax, so two representations should count as "close" when they produce similar distributions. That leads them to information geometry: the log-normalizer A plays the role of a convex potential, its gradient φ(λ) = ∇A(λ) gives a dual coordinate equal to the expected item embedding, and straight lines in the primal versus dual coordinate systems give two different interpolations. They analyze what each interpolation does to the output distribution, showing one behaves like an intersection and the other like a union.

They then formalize steering as a constrained optimization: move the representation onto the hyperplane where the probe score equals a target value c, while minimizing some distance from the original representation. With Euclidean distance the solution is the standard method of adding the probe vector; with the KL divergence the solution turns out to add the probe vector in the dual space instead, which is dual steering. They prove this under a concept-factorizable assumption, where the output space splits into counterfactual pairs sharing a semantic attribute plus neutral items.

For implementation, they cannot move directly in dual space because φ must stay inside the convex hull of item embeddings, so they trace the corresponding curved path in primal space using a first-order Taylor expansion of ∇A, which involves the covariance of embeddings under the softmax distribution. Because that covariance is often low-rank, they add a regularization term α I_d and solve the system with conjugate gradients, iterating small steps to trace the path. Experiments use Gemma-3-4B with contexts from AllenAI C4 for language concepts and MetaCLIP-2 with synthetic object images and COCO for vision concepts, comparing Euclidean steering against dual steering along both Primal MD and Dual MD probe directions.

Why This Matters

This work shifts the discussion of the linear representation hypothesis from "which direction encodes the concept" to "what geometry the representation space actually has," and it supplies a concrete, provably motivated method plus a practical algorithm and code. It suggests that many brittleness problems in interpretability are geometric mismatches rather than failures of the linear hypothesis itself, and that similar reasoning could be applied beyond softmax models.

Real-world applications:

  • Safer language model deployment: reliably shifting attributes such as verb tense or language output while keeping unrelated predictions stable.
  • Multimodal retrieval and generation with models like CLIP: changing a target object (for instance dog to cat) without accidentally promoting mixtures such as "cat + dog" images.
  • Controlling sensitive attributes in generated text, such as gender-related token choices, without perturbing unrelated vocabulary.
  • Serving as a template for geometry-aware interventions in other model families where outputs are produced through a softmax-style distribution.

Industry relevance: teams that build guardrails, personalization, or content-moderation interventions on top of large models can use dual steering to make edits that are more controllable and less damaging to surrounding behavior, with a documented complexity improvement (O(nkd)) that makes iterative implementation practical at vocabulary scale. The released code at github.com/KihoPark/dual-steering lowers the barrier to adoption.

Future Directions

  • Extending the information-geometric framework beyond softmax-based models to other architectures and output distributions, as the authors suggest the high-level idea is widely applicable.
  • Understanding what makes an ideal linear probe and how to identify one, since steering quality depends heavily on probe quality (the paper explicitly assumes a probe has been identified and notes its analysis is truncated in the provided content at the probing-assumption discussion).
  • Relaxing or testing the concept-factorizable distribution assumption, which the authors note may not always hold in practice even though the Euclidean/dual difference does not require it.
  • Refining the regularized Newton procedure, including how the regularization parameter α and step schedule interact with low-rank Hessians when the distribution is highly concentrated.

Target Audience

Advanced machine learning researchers and graduate students working on interpretability, representation geometry, or model steering, along with practitioners who already use linear probes and activation steering and want a more theoretically grounded, robust alternative. Readers need comfort with convex analysis and KL divergence; the paper's intuitions are explained accessibly, but the proofs and derivations assume a strong mathematical background.

Authors’ abstract

This paper concerns the question of how AI systems encode semantic structure into the geometric structure of their representation spaces. The motivating observation of this paper is that the natural geometry of these representation spaces should reflect the way models use representations to produce behavior. We focus on the important special case of representations that define softmax distributions. In this case, we argue that the natural geometry is information geometry. Our focus is on the role of information geometry on semantic encoding and the linear representation hypothesis. As an illustrative application, we develop "dual steering", a method for robustly steering representations to exhibit a particular concept using linear probes. We prove that dual steering optimally modifies the target concept while minimizing changes to off-target concepts. Empirically, we find that dual steering enhances the controllability and stability of concept manipulation.

Read the original paper