Skip to content
AI.info

Research

Generative Late-Interaction Embeddings For Visual Document Retrieval

Overview Research area: Information retrieval (cs.IR), specifically late-interaction multi-vector retrieval over visual documents (screenshots of pages containing tables, figures, charts, and text). T

Generative Late-Interaction Embeddings For Visual Document Retrieval
arXiv
2609.11808
Published
2026-09-10
Authors
Mohamed Eltahir, Talal Aloushan, Rose Khairoalsendi, Jana Shata, Mohammed Alhassan, Leen Alrehaili, Tanveer Hussain, Naeemullah Khan

AI summary

Overview

  • Research area: Information retrieval (cs.IR), specifically late-interaction multi-vector retrieval over visual documents (screenshots of pages containing tables, figures, charts, and text).
  • Technical level: Advanced. The paper assumes familiarity with MaxSim scoring, ColBERT-style late interaction, k-means clustering, and neural decoders, and it uses geometric arguments (intrinsic dimension estimation, unit-sphere constraints) as the backbone of its design.
  • Scope: The paper measures the geometry of stored page embeddings produced by frozen visual-document encoders, then builds a post-hoc codec that stores only a handful of vectors per page and regenerates the full vector set on demand for exact rescoring.

What This Paper Is About

Late-interaction retrieval is the strongest known approach for searching visual documents, but it stores roughly a thousand vectors per page (ColPali: 1,031 patch vectors of dimension 128, about 258 KB in bfloat16, so one million pages cost about a quarter terabyte of embeddings). Existing compression methods keep a subset or a local average of those vectors and all stop at roughly sixteen vectors per page, while the methods that go below that floor require retraining the encoder and re-encoding the entire corpus. This paper asks what the stored set actually is geometrically, finds that each page's vectors lie on the unit sphere and concentrate near a manifold of intrinsic dimension five to six, and uses that finding to build a method that stores a few vectors per page and regenerates the rest only for the top-ranked candidates.

Key Contributions

  1. Geometry of the stored object. The authors measure page token clouds across three encoders and all ten ViDoRe v1 corpora (6,729 pages) and report that the vectors sit exactly on the unit sphere and have a median TwoNN intrinsic dimension of 4.9 on ColPali (5.1 on ColQwen2, 6.1 on Nemotron v2 at an ambient dimension of 3,072).
  2. Spherical anchoring. They show that standard k-means centroids of unit vectors fall inside the sphere, which systematically understates MaxSim scores, and that re-projecting centroids onto the sphere is a free correction worth up to +0.093 nDCG@5 over raw centroids.
  3. Generative Late-Interaction Embeddings (GLIE). They replace extractive subset sampling with a codec that stores k vectors per page and regenerates all N from them on demand, with a zero-initialized refiner that provably starts at normalized clustering and a decoder that structurally cannot score worse than the stored code.
  4. An asymmetric two-stage pipeline. Initial retrieval runs on the k stored vectors alone, and only the top-L shortlist (L = 20) is expanded back to N vectors and rescored exactly, making the expensive step apply to almost nothing.

Main Findings

  • Aggressive storage budgets work. GLIE reaches 79% of the uncompressed system's nDCG@5 (0.657 against a ceiling of 0.836) on ViDoRe v1 at k = 4, which the paper describes as nearly 80% versus 70% for the best prior post-hoc method, and 91% at k = 16 (about 4 KB per page). It beats every prior baseline on all subsets of both ViDoRe v1 and v2 at every budget.
  • The geometry is consistent across encoders and corpora. Per-corpus medians on ColPali span 4.7 to 5.1, from scientific figures to government reports. Ambient dimension varies by 24 times across the three encoders while intrinsic dimension varies by one. A Gaussian fitted to each page's own covariance reads 32.2 on the same pages, and uniform noise of the same size reads 61.4, so the low dimension is a property of the token cloud rather than of the estimator.
  • Normalizing centroids is the single largest free win. Spherical anchoring contributes from +0.030 at k = 64 up to +0.093 at k = 4, and the paper states it costs nothing. The improvement shrinks as k grows, exactly as the proposition relating centroid norm to cluster spread predicts.
  • The learned code adds less than the free correction. The refined code adds +0.044 at k = 2 down to +0.016 at k = 16 and nothing beyond that budget, while the generative read-out is the aggressive-budget specialist, peaking at +0.016 at k = 4.
  • Margins concentrate at small budgets on v1 but not on v2. On ViDoRe v1 the margin forms a plateau of about 0.04 across k ≤ 8, halves at k = 16, and decays to noise by k = 32, a mean of +0.039 below k = 8 against +0.010 above. The stored code alone turns slightly negative at k = 64, where normalized clustering already sits 0.027 from the ceiling, and four of the ten v1 subsets have ceilings of 0.94 to 0.98. On ViDoRe v2, which the paper says saturates nowhere, the margin never decays out to k = 64.
  • Decoding, not the shortlist, is the binding constraint. At k = 4 on v1, GLIE scores 0.657, the shortlist oracle scores 0.782, and the uncompressed ceiling is 0.836; the first gap is decode fidelity and the second is shortlist recall. Widening the shortlist from L = 5 to L = 100 raises the oracle from 0.705 to 0.822 while GLIE moves only from 0.647 to 0.660, and widening L from 20 to 100 is worth only +0.002.
  • Training budget efficiency is extreme. The full system uses a 415K-parameter network fitted in under three GPU-minutes on roughly a thousand training pages. Fine-tuning the encoder on 4,000 pages with 13.3M LoRA parameters over 1.5 GPU-hours does not reach even the training-free stage of GLIE at any budget, and GLIE beats it at all six budgets by +0.074 to +0.132.
  • The recipe transfers to a second encoder. On ColQwen2, GLIE retains 82% of uncompressed quality at k = 4 and improves all ten subsets at every budget from k = 4 up, losing only at k = 2, where it improves 7 of 10 subsets.
  • Hyperparameters are not the source of the gains. Varying decoder width from 128 to 1024 and depth from one block to three, a seventy-fold range from 184K to 13M parameters, moves nDCG@5 by at most 0.009 with no monotone trend, and the smallest decoder trained is the best at k = 4.
  • Data efficiency. Refitting on 1,250, 2,500, and 5,000 source pages leaves margins statistically unchanged (+0.048, +0.052, +0.049 at k = 4), so the codec saturates on roughly a thousand pages.
  • Storage figures. At k = 4 the stored representation is 1,040 bytes per page against 257.8 KB uncompressed, so one million pages shrink from 258 GB to 1.0 GB.

Methodology in Plain English

The authors begin by measuring rather than designing. They take the token vectors that a frozen encoder produces for a page and estimate the intrinsic dimension of the cloud, finding that a page with 1,031 vectors of dimension 128 behaves like a surface with only five or six degrees of freedom, and that every vector sits exactly on the unit sphere because the encoders L2-normalize their output.

This geometry produces two design decisions. First, since a k-means centroid is a Euclidean average of unit vectors, it lies inside the sphere and therefore understates every inner product used by MaxSim. Projecting each centroid back onto the sphere fixes that at no cost, and this becomes the training-free floor for the whole system. Second, since a page has so few degrees of freedom, it should be possible to describe the page with a few vectors and regenerate the rest from that description instead of permanently discarding information.

The codec is built so that training can only improve on the training-free baseline. A refiner, a cross-attention layer with a zero-initialized output projection, adjusts the projected centroids while reading the full token set, so at initialization it reproduces normalized clustering exactly. A decoder then expands the k stored vectors back into N vectors for reranking, structured with count-proportional slots so each cluster gets as many outputs as it has real patches, an exact anchor so each cluster's first output is the stored vector itself, and a bounded step capped at a factor of 0.75 so generated children stay near their anchor. Because MaxSim is a maximum and the decoded set contains the stored set, regeneration can add evidence but cannot lower a stored score.

Training uses five kinds of loss. Two terms, a per-query-token MaxSim match and a listwise KL over the candidate ranking, are applied to both the code and the regenerated set so each retrieves like the full page. Three terms apply only to the regenerated set: a one-sided penalty that charges only when a non-relevant candidate is scored above the teacher, a Chamfer distance within each cluster so generated vectors occupy the region the real patches occupy, and a support-function match over 128 fixed random directions so the generated set reaches as far as the real page in every direction. Reconstruction error is deliberately excluded, since minimizing it would collapse every child toward its cluster mean and lose the extreme points that MaxSim reads.

At query time, stage one scores every page by MaxSim over its k stored vectors, and stage two expands only the top 20 candidates back to N vectors for exact rescoring; everything outside the shortlist keeps its first-stage order. The encoder is never updated, so the entire fit happens post hoc on cached embeddings.

Why This Matters

The paper reframes index compression for late-interaction retrieval as a question about what is being stored rather than how much of it to keep, and it identifies a free correction (normalizing cluster centroids) that any dot-product late-interaction system can adopt today. It also demonstrates that a frozen-backbone, 415K-parameter codec fitted in under three GPU-minutes on a thousand pages can outperform a fine-tuned encoder at a matched budget, which changes the economics of building compressed visual-document indexes.

Real-world applications:

  • Enterprise and legal document search over scanned or visually rich pages, where a million-page corpus currently costs 258 GB of embeddings and would cost 1.0 GB at k = 4.
  • Retrieval-augmented generation over financial filings, scientific figures, and government reports, where local evidence such as a table cell or a caption phrase must survive compression.
  • On-device or edge deployment of visual document search, where a 258 GB index is infeasible but a 1 GB index with on-demand decoding is not.
  • Video retrieval, where token sets are largest and most redundant and where the paper expects the largest gains.

Industry relevance centers on cost: the storage saving is roughly 250-fold at k = 4, the fit takes GPU-minutes rather than GPU-hours, and changing the storage budget of a deployed index touches only cached embeddings rather than requiring the corpus to be re-encoded. That last property matters for vendors who cannot give up a public checkpoint or re-run encoding pipelines over customer corpora.

Future Directions

  • Better decoders. The authors state this is the immediate move on their axis, since perfectly decoding the same code and shortlist at four vectors per page would reach 0.782 against their 0.657, and a further +0.13 sits in the shortlist waiting on improved decoding.
  • Extending the sweep to further late-interaction encoders. Two encoders are evaluated (ColPali v1.3 and ColQwen2) plus geometry measurements on a third (Nemotron v2); how far the recipe carries beyond these is untested.
  • Composing GLIE with quantized storage. The paper frames quantization and vector-count compression as acting on different axes of the footprint, so the two should combine.
  • Applying the method where token sets are largest and most redundant, which the authors identify as video late interaction.

Target Audience

This paper is for retrieval researchers and engineers working on multi-vector or late-interaction systems, particularly those maintaining visual-document indexes under storage pressure. It is also relevant to practitioners who cannot retrain encoders because they depend on public checkpoints or already-cached embeddings, and to readers interested in empirical geometry of learned representations, since the intrinsic-dimension measurements and the centroid-norm proposition stand on their own apart from the retrieval system built on them.

Authors’ abstract

Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means centroids fall inside the sphere, causing systematic underestimation of MaxSim scores. Normalizing them to the surface is a free correction worth up to +0.093 nDCG@5 over raw centroids. Second, because the page manifold has few degrees of freedom, the full set of vectors can be regenerated from only a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): k << N vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating the page's full embedding set. At query time, search runs exclusively on these k vectors, and a decoder expands only the top candidates back to all N vectors for exact rescoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in under three GPU-minutes on just a thousand training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns hold across a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.

Read the original paper