Research
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings Overview Research area: Multimodal representation learning and mechanistic interpretability, specifically the geometry of shared v

- arXiv
- 2602.06218
- Published
- 2026-02-05
- Authors
- Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard
AI summary
Cross-Modal Redundancy and the Geometry of Vision–Language EmbeddingsOverview
Research area: Multimodal representation learning and mechanistic interpretability, specifically the geometry of shared vision–language embedding spaces probed with sparse autoencoders (SAEs). Posted to arXiv (2602.06218v2 [cs.CV], 09 Feb 2026) by Grégoire Dhimoïla (Brown University, ENS Paris Saclay, IRT Saint Exupéry), Thomas Fel (Kempner Institute, Harvard), Victor Boutin (CNRS) and Agustin Picard (IRT Saint Exupéry).
Technical level: Advanced. The paper assumes familiarity with sparse coding, dictionary learning, nonlinear ICA identifiability arguments, contrastive dual encoders, and the modality-gap literature.
Scope: The paper introduces a testable inductive bias (Iso-Energy) for recovering which concepts a vision–language model represents across both modalities, and uses it to decompose, interpret and intervene on the embedding geometry of six foundation-scale VLMs.
What This Paper Is About
Vision–language models such as CLIP and SigLIP place images and text into a shared embedding space, but exactly how that space organizes cross-modal concepts is poorly understood. Sparse autoencoders trained on these embeddings tend to produce dictionaries that split cleanly by modality, mirroring the well-documented "modality gap," and standard reconstruction metrics cannot distinguish a dictionary that captures shared concepts from one that does not. The paper's goal is to supply a principled criterion for identifying genuinely shared (bimodal) concepts and to show that the resulting decomposition is both interpretable and directly actionable.
Key Contributions
- The Iso-Energy Assumption. A concept that is genuinely shared across modalities should exhibit the same average energy (average squared activation) in image and text inputs. This gives a concrete, domain-agnostic criterion for labeling an atom as bimodal versus unimodal.
- The Aligned Sparse Autoencoder (SAE-A). An alignment-penalized Matching Pursuit SAE that enforces energy consistency during training while preserving reconstruction, validated against ground-truth synthetic data on which a classical SAE fails.
- A geometric decomposition of foundation-scale VLMs. Sparse bimodal atoms carry the entire cross-modal alignment signal, while unimodal atoms carry modality-specific information and fully explain the modality gap, with a few high-energy atoms acting as modality-specific biases. The authors state this contrasts with idiosyncratic atoms described by Papadimitriou et al. (2025).
- Actionable interventions. Removing unimodal atoms collapses the modality gap without harming retrieval or zero-shot performance, and restricting vector arithmetic to the bimodal subspace yields in-distribution queries and improved retrieval.
Main Findings
- Controlled sanity checks pass. On synthetic data matching CLIP-like cosine similarity statistics (sparse code with ‖z‖₀ = L = 20), when Iso-Energy is violated (τ₁ ≠ 1) both SAE and SAE-A recover the dictionary equally well (𝒲 ≈ 0.19, mma ≈ 0.82), showing the regularizer does not hallucinate bimodal atoms. When Iso-Energy holds (τ₁ = 1), the standard SAE fails (𝒲 = 0.396, mma = 0.29) while SAE-A succeeds (𝒲 = 0.184, mma = 0.52). The bias is therefore neutral when unnecessary and decisive when appropriate.
- Reconstruction is essentially unchanged. Across the six VLMs the paper reports MSE values between 0.115 and 0.257 and R² between 0.742 and 0.885; the paper describes the two variants as nearly identical, and synthetic experiments showed R² ≥ 0.99.
- Alignment is carried entirely by bimodal atoms. Functional alignment ρ increases in every (SAE, SAE-A) pair, with aligned values ranging from 1.475 to 16.58 versus 0.327 to 8.737 for the unregularized SAE — the paper describes this as an increase of more than an order of magnitude. FDA (Functional and Distributional Agreement) increases in every pair, from a range of 2.630–9.787 for SAE to 4.559–34.95 for SAE-A, which the paper characterizes as doubling or tripling.
- Unimodal atoms are safe to ablate. Interventional robustness δ_r stays small: the largest unregularized value is 0.224, and aligned values range from −0.006 to 0.125. Since δ_r measures the change in recall after ablating unimodal atoms, a small value means these atoms contain no cross-modal information.
- Energy distribution is highly skewed. Most features are bimodal and medium-energy, while a handful of high-energy unimodal features dominate modality-specific variance and behave as modality biases.
- Three-cluster geometry. Low-dimensional projections of the learned atoms separate into image-only, text-only and bimodal clusters. Unimodal atoms align with the modality cones of the embedding space; bimodal atoms occupy a modality-agnostic subspace orthogonal to those directions, explaining why each type carries the information it does.
- Semantic inspection is consistent. Bimodal atoms are semantically stable across modalities (colors, objects, actions), whereas unimodal atoms typically reflect idiosyncratic signals such as image cropping artifacts or text "name patterns."
- Closing the modality gap. Masking out unimodal atoms collapses the gap and merges the image and text distributions, unlike the embedding-shift baseline of Liang et al. (2022), which matches means but leaves distributions separated. The paper measures the gap using an out-of-distribution method borrowed from Sun et al. (2022), based on the distance from each image (in-distribution) and caption (out-of-distribution) embedding to its 10th nearest image neighbor.
- In-distribution semantic arithmetic. Restricting edits to bimodal atoms reduces the OOD score in every model tested: CLIP 0.97 → 0.77, CLIP-L 0.95 → 0.76, OpenCLIP 0.86 → 0.68, OpenCLIP-L 0.87 → 0.72, SigLIP 0.99 → 0.70, SigLIP2 0.99 → 0.61. Performance is preserved on the FashionIQ benchmark (Wu et al., 2021).
Methodology in Plain English
The authors model each image or caption as the output of a generative process: a sparse vector of latent concepts is sampled, then rendered through a domain-specific generator. A VLM encoder is treated as a partial inverse of this process, and a sparse autoencoder tries to undo the encoder and recover the concepts. Because recovering concepts from nonlinear generators is provably unidentifiable without extra constraints, the authors add one: a soft penalty that encourages the same atom to have similar activation strength on image and text inputs.
Concretely, they take a Matching Pursuit SAE (based on Costa et al., 2025, and Mallat and Zhang, 1993) as the base model and add an alignment term equal to the negative normalized trace of the product of image and text code matrices. They first validate the recipe on synthetic data where ground truth is known, then train both the standard SAE and the aligned SAE-A on activations from six representative models — CLIP (ViT-B/32, ViT-L/14), OpenCLIP, OpenCLIP-L, SigLIP and SigLIP2 — using identical hyperparameters (expansion ratio 8, target ℓ₀ = 20) and a random subset of 1 million LAION embeddings, so that differences are attributable to the regularizer. The regularization weight β is chosen automatically by a log-sweep over {10⁻⁶, …, 10⁻¹}, taking the largest value whose explained variance differs from the unregularized SAE by less than 0.05; the working value is approximately 10⁻⁴.
Because reconstruction metrics cannot distinguish concept structures, the authors introduce four multimodality-sensitive metrics: probing accuracy p_acc (how well atoms predict domain membership), functional alignment ρ (whether bimodal features drive alignment), FDA (distributional agreement), and interventional robustness δ_r (change in recall after ablating unimodal atoms). They then test interventions, including re-expressing activations as (Z ⊙ δ)D with a binary mask that filters unimodal atoms, and comparing against the mean-shift baseline of Liang et al. (2022).
Why This Matters
Impact on research. The paper argues that interpretability of multimodal models has been limited by dictionaries that mix shared and modality-specific structure without distinguishable reconstruction quality. Iso-Energy gives a falsifiable criterion for separating the two, provides a concept-level account of the modality gap that complements geometric (conical) and training-dynamics explanations, and offers Proposition 1 showing that removing modality information preserves ranking when visual and textual information are orthogonal — a generalization of Zhang et al. (2023). The authors also show that interventions derived purely from canonical basis directions (Schrodi et al., 2025) do not describe the modality-specific subspace while damaging the shared one.
Real-world applications (drawn from the application areas the paper cites for VLMs):
- Visual question answering systems (Chen et al., 2024) that must reason jointly over image and text content.
- Medical imaging pipelines (Singhal et al., 2023) where a modality-gap artifact could distort image–report matching.
- Autonomous driving (Zhou et al., 2024), where vision–language embeddings support scene understanding and instruction following.
- Embodied AI (Shukor et al., 2025) and image retrieval, where the paper's bimodal-restricted arithmetic produces queries that stay on the semantic manifold, as demonstrated on FashionIQ.
Industry relevance. The technique is a drop-in diagnostic and intervention layer for any dual-encoder retrieval or grounding stack: it can close the modality gap without the retrieval degradation associated with mean-shift fixes, and it supports controlled semantic edits that remain in-distribution. Code is released at https://github.com/Parabrele/IsoEnergy.
Future Directions
- Robustness of the weighting coefficient. The paper's own limitations text states that the alignment penalty is sensitive to the choice of β; the provided content is truncated at this point, so the remaining limitations described by the authors are not reported here.
- Scaling beyond two domains. The generative formalism is stated for an arbitrary set of domains 𝔇, but the experiments cover image and text only; extending Iso-Energy to audio, video or other modalities is left open.
- Alternative dictionary learners. Iso-Energy is operationalized on a Matching Pursuit SAE; whether ReLU, JumpReLU or BatchTopK SAEs would show the same bimodal/unimodal separation is not tested.
- Downstream validation of interventions. The paper evaluates gap closure and OOD behavior plus retrieval on FashionIQ; whether removing unimodal atoms preserves performance on broader cross-modal and unimodal downstream benchmarks is not reported.
Target Audience
Authors’ abstract
Vision-language models (VLMs) align images and text with remarkable success, yet the geometry of their shared embedding space remains poorly understood. To probe this geometry, we begin from the Iso-Energy Assumption, which exploits cross-modal redundancy: a concept that is truly shared should exhibit the same average energy across modalities. We operationalize this assumption with an Aligned Sparse Autoencoder (SAE) that encourages energy consistency during training while preserving reconstruction. We find that this inductive bias changes the SAE solution without harming reconstruction, giving us a representation that serves as a tool for geometric analysis. Sanity checks on controlled data with known ground truth confirm that alignment improves when Iso-Energy holds and remains neutral when it does not. Applied to foundational VLMs, our framework reveals a clear structure with practical consequences: (i) sparse bimodal atoms carry the entire cross-modal alignment signal; (ii) unimodal atoms act as modality-specific biases and fully explain the modality gap; (iii) removing unimodal atoms collapses the gap without harming performance; (iv) restricting vector arithmetic to the bimodal subspace yields in-distribution edits and improved retrieval. These findings suggest that the right inductive bias can both preserve model fidelity and render the latent geometry interpretable and actionable.