Skip to content
AI.info

Research

Learning Multimodal Embeddings with Evidence-Aligned Readout

Learning Multimodal Embeddings with Evidence-Aligned Readout Overview Research area: Multimodal retrieval and representation learning — specifically, how multimodal large language models (MLLMs) can t

Learning Multimodal Embeddings with Evidence-Aligned Readout
arXiv
2609.33659
Published
2026-09-27
Authors
Zirong Chen, Fuda Ye, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Jiachuan Wang, Yongqi Zhang

AI summary

Learning Multimodal Embeddings with Evidence-Aligned Readout

Overview

Research area: Multimodal retrieval and representation learning — specifically, how multimodal large language models (MLLMs) can turn generated, task-relevant "evidence" into a single retrieval embedding.

Technical level: Advanced. The paper assumes familiarity with contrastive retrieval training, teacher forcing, MLLM hidden states, and factorial experimental design.

Scope: The paper proposes EviAlign, a framework that couples Semantic Evidence Generation with Boundary Readout in one shared MLLM, and tests through matched controls whether the semantic organization of generated evidence should determine where retrieval representations are read.

What This Paper Is About

Multimodal embedding models must compress text, images, and interleaved inputs into a shared space so that relevant matches can be found. MLLMs can already generate task-relevant cues (the authors call these "retrieval evidence"), but producing useful evidence does not by itself determine how that evidence enters the final retrieval embedding. This paper asks a narrower design question — whether the semantic units produced during generation can also specify where the model reads its internal representations — and answers it with a controlled 2×3 study plus a full model trained on 500K query–candidate pairs.

Key Contributions

  1. Formulating evidence-aligned readout. The paper defines a co-design principle with three requirements: generation should select task-relevant evidence, the organization of that evidence should explicitly determine the semantic roles and locations of readouts, and the resulting readouts should be jointly integrated into a retrieval-optimized embedding.

  2. Introducing EviAlign. A concrete instantiation that organizes evidence into five semantic units — Entity, Attribute, Relation, Detail, and Summary — with boundary tokens <ENT>, <ATT>, <REL>, <DET>, and <SUM> (new vocabulary items initialized from the [EOS] embedding), then mean-pools the five contextualized boundary states into one L2-normalized vector, preserving single-vector indexing and scoring.

  3. A matched 2×3 controlled study. Comparing consistent versus permuted evidence organization across trailing, distributed, and boundary readouts, with training targets containing the same five evidence spans — isolating the interaction between organization and readout location.

  4. A decomposition of generation, readout capacity, and evidence source. Showing through controlled variants that neither free-form chain-of-thought alone nor additional readout tokens alone reproduce the boundary-readout result.

Main Findings

  • Strong retrieval results. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs on Qwen3-VL-8B, and 75.5 on Qwen2-VL-7B. The Qwen3-VL-8B model reaches 81.6 in-domain and 67.7 out-of-domain.

  • A 1.74-point co-design interaction. In the 2×3 study, the Semantic–Mixed advantage is +0.14 under Trailing, +0.65 under Distributed, and +2.39 under Boundary. Equivalently, moving from distributed positions to boundaries adds +0.35 for Mixed and +2.09 for Semantic. Both comparisons give Δco = 2.39 − 0.65 = 2.09 − 0.35 = 1.74 points.

  • Absolute six-configuration results. Mixed: 73.60 (Trailing), 74.20 (Distributed), 74.55 (Boundary). Semantic: 73.74 (Trailing), 74.85 (Distributed), 76.94 (Boundary).

  • Semantic evidence and free-form CoT are nearly tied under a trailing readout. CoT plus a single trailing readout reaches 73.75; semantic evidence with a single trailing readout reaches 73.74.

  • Readout multiplicity alone is not enough. Five readouts without generation reach 70.30; five trailing readouts after semantic evidence reach 73.07; the boundary-readout model reaches 76.94, 3.20 points higher than the separately trained trailing-readout model.

  • Evidence source matters less than the boundary interface. Changing the evidence teacher from GLM-4.1V-9B-Thinking to Qwen3-VL-8B-Thinking gives 75.60 versus 76.94. Prepending free-form CoT to semantic evidence gives 76.59 with the same boundary readout.

  • Boundary readouts retain complementary information. ENT alone reaches 76.12; all five reach 76.94; removing any one reduces the average by 0.23–0.36 points. Detail is the strongest single readout on CIRR (78.6), Attribute on VisualNews image-to-text (84.8), and Relation on OVEN (70.9). On OVEN, the full aggregate reaches 72.2. Full aggregation exceeds every individual readout on seven tasks.

  • Five units is a sweet spot. Training K ∈ {1, 3, 5, 7} evidence-schema/readout configurations gives 72.8 (K=1), 75.4 (K=3), 76.9 (K=5), and 76.2 (K=7). The extra units for K=7 are Context and Emotion.

  • Backbone scaling helps most out of domain. Scaling from 2B to 7B/8B improves overall results from 72.6 to 75.5 for Qwen2-VL and 73.4 to 76.9 for Qwen3-VL. Out-of-domain gains are 5.8 and 8.2 points, versus in-domain gains of 1.4 and 1.2 points.

  • Format validity is high but generation is costly. Of 12,000 query generations, 99.52% contain all five boundary tokens exactly once and in order. At batch size 4, end-to-end batch times are 94.6 ms for direct encoding, 1,502.5 ms for EviAlign, and 9,825.3 ms for CoT-plus-semantic Boundary Readout — the latter two averaging 114.3 and 702.1 generated tokens and reaching 76.94 and 76.59 Recall@1 respectively.

Methodology in Plain English

The setup starts from a simple interface: given an input (text, image, or interleaved input plus a task instruction), a single MLLM both generates text and produces one embedding vector.

Generation side. Instead of free-form captions or chain-of-thought, the model is trained to emit five labeled evidence units in a fixed order, each ending with a special marker token. Entity captures objects or concepts with identifying properties; Attribute captures scene-level characteristics; Relation describes actions, interactions, and spatial relations; Detail keeps locally discriminative cues such as visible text and fine-grained patterns; and Summary gives a retrieval-focused global description.

Readout side. Because the marking tokens sit at the end of each unit, their positions are semantically meaningful. Under causal attention, the hidden state at <ENT> has seen the input and the Entity unit, the state at <ATT> has seen all of that plus Attribute, and so on. The model extracts these five final-layer hidden states, mean-pools them, and L2-normalizes the result into the single vector used for scoring.

Training. Each input is paired with a target evidence sequence. Under teacher forcing, boundary readouts are taken at the boundary token positions in that target rather than in a generated sequence. Two losses are summed: a symmetric in-batch contrastive loss (query-to-candidate and candidate-to-query) that shapes the pooled embedding, and a language-modeling loss that teaches the model to produce the evidence in the prescribed structure. The vision encoder is frozen; the LLM and projector are fine-tuned for one epoch on 500K query–candidate pairs from the MMEB-V1 retrieval training split. Evidence targets were generated by GLM-4.1V-9B-Thinking.

Inference. No target is given. The model greedily generates its own evidence sequence, reads hidden states at the generated boundaries, and pools them. If a boundary token is missing, its readout falls back to the final-layer hidden state of the last generated token. Queries and candidates are encoded independently with their own instructions, so candidate embeddings can be precomputed and indexed. Retrieval uses plain cosine similarity.

The controlled comparison. To isolate whether evidence organization matters, the researchers built Semantic targets (fixed correspondence between evidence roles and boundary positions) and Mixed targets (the same five labeled spans, permuted per example, with boundary-token identities and order unchanged). They crossed this with three readouts: Trailing (only the final <SUM> state), Distributed (five readout tokens at positions ⌊kT/5⌋, where T is the evidence-target length before insertion), and Boundary (the same five token types at the five evidence-unit ends). All six configurations were trained independently under the same budget, data, and optimization schedule.

Why This Matters

Impact on research. The paper reframes a question that generation-enhanced retrieval systems tend to leave implicit: generating good evidence and using it well are separate problems. Its central result — that semantic evidence and free-form CoT are nearly tied under a trailing readout (73.74 vs. 73.75) but diverge sharply at boundaries — suggests that some reported gains from "reasoning before embedding" may be attributable to readout geometry rather than content. The 2×3 factorial design with matched evidence spans is a reusable template for other co-design claims.

Real-world applications (grounded in the evaluated task suite):

  • Composed image retrieval for shopping. A user supplies a reference photo plus a modification ("show three bottles," "is darker blue and has a V-neck"); the system must find the modified target. This is the CIRR and FashionIQ setting, where Detail is the strongest single readout on CIRR (78.6) and Attribute on VisualNews i2t (84.8).
  • News and media asset search. VisualNews t2i and i2t test entity-rich caption-to-image and image-to-caption retrieval at scale.
  • Knowledge-intensive evidence retrieval. WebQA isolates the evidence-retrieval stage of multimodal question answering; OVEN requires entity grounding (Relation is its strongest single readout at 70.9, with the full aggregate at 72.2).
  • Visual document understanding. The benchmark suite includes visually rich document retrieval, where locally discriminative cues like visible text matter.

Industry relevance. The design keeps the deployment interface unchanged: aggregation into one normalized vector means standard single-vector indexing and cosine scoring still work, and candidate embeddings can be precomputed offline. That matters for systems built on existing vector databases. The cost side is also explicit — EviAlign's batch time at batch size 4 is 1,502.5 ms versus 94.6 ms for direct encoding, and adding a CoT prefix pushes it to 9,825.3 ms without improving quality. Query-side generation remains online; only candidate encoding is offline.

Future Directions

  1. Reducing generation cost. EviAlign's online batch time is roughly 16× direct encoding at batch size 4, and adding CoT raises it to 9,825.3 ms while slightly lowering Recall@1 (76.59 vs. 76.94). The paper does not explore ways to shorten the evidence sequence without losing the boundary structure.

  2. Understanding why more units stop helping. Performance rises from 72.8 at K=1 to 76.9 at K=5, then falls to 76.2 at K=7 with added Context and Emotion units. The paper reports the pattern but offers no mechanism for the decline, leaving open how to design schemas for new domains.

  3. Closing the in-domain/out-of-domain gap. Even the best configuration reports 81.6 in-domain against 67.7 out-of-domain, and GME records the strongest out-of-domain average at 71.8. Boundary readouts help out-of-domain most under scaling (gains of 5.8 and 8.2 points), but the absolute gap remains large.

  4. Reducing dependence on a teacher model. Evidence targets were produced by GLM-4.1V-9B-Thinking, and swapping teachers changes the result (75.60 for Qwen3-VL-8B-Thinking vs. 76.94 for GLM). Whether the schema can be induced without a large external annotator is not addressed.

  5. Comparison beyond single-vector retrieval. The paper positions EviAlign against multi-vector late-interaction systems such as ColBERT, ColPali, and MetaEmbed by design choice rather than direct experiment; the analysis section does not report a head-to-head comparison on the same 12-task protocol.

Target Audience

Researchers and engineers working on multimodal retrieval, embedding models, and MLLM-based representation learning — particularly those building generation-enhanced or reasoning-based embedders who need to distinguish gains from generated content versus gains from how that content is read out. It also suits practitioners evaluating whether to add query-side generation to an existing single-vector retrieval index, since the paper reports both the accuracy and the latency of that trade-off. Readers should be comfortable with contrastive training, teacher forcing, and factorial experimental design.

Authors’ abstract

Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled $2\times3$ study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.

Read the original paper