Skip to content
AI.info

Research

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers

Overview Research area: Information retrieval (cs.IR), specifically dense retrieval over mixed-modality corpora, sitting at the intersection of multimodal representation learning (CLIP-based and visio

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
arXiv
2610.11816
Published
2026-10-08
Authors
Yubo Sun, Chunyi Peng, Yukun Yan, Zhenghao Liu, Zhipeng Xu, Sen Mei, Linlin Xin, Zheni Zeng, Maosong Sun

AI summary

Overview

Research area: Information retrieval (cs.IR), specifically dense retrieval over mixed-modality corpora, sitting at the intersection of multimodal representation learning (CLIP-based and vision-language model architectures) and retrieval evaluation.

Technical level: Intermediate. The core idea is accessible, but the paper assumes familiarity with dense retrievers, contrastive objectives such as InfoNCE, and multimodal encoders.

Scope: The paper diagnoses a systematic bias in how mixed-modality retrievers score text versus image documents, then proposes a training method intended to correct it. The abstract reports findings and a proposed solution; it contains no numbers, so no quantitative results can be reported here.

What This Paper Is About

Dense retrievers work well on text-only and image-only collections, but real collections are usually mixed — containing text documents, image documents, and documents that fuse the two. This paper asks whether retrieval quality holds up when those modalities coexist, and finds that it often does not. The goal is to explain why mixed-modality retrieval degrades and to design a training approach that removes the underlying bias.

Key Contributions

  1. A systematic sweep across architectures and modality compositions. The authors examine retrievers across multiple architectures and vary the makeup of the corpus, rather than testing a single model on a single benchmark.
  2. Identification of a V-shaped performance curve. As image documents are progressively replaced by semantically corresponding text representations, retrieval performance drops in the middle (mixed corpora) and recovers at the ends (single-modality corpora).
  3. Naming and analyzing "Chaos in the Text." The paper reports that irrelevant text degrades retrieval more severely than an equal number of irrelevant images, and traces this to a modality preference in the similarity space — text representations receive systematically higher similarity scores, so irrelevant text can outrank relevant images.
  4. The Trident method. A training approach that builds text, image, and fused text-image views of each document as co-equal positives, optimized with Multi-Positive View InfoNCE to jointly handle relevance discrimination and balance across positive views.

Main Findings

  • Mixed corpora are the weak point. Performance stays strong when a corpus is purely one modality but degrades substantially when modalities coexist in the same index, producing the V-shaped curve described in the abstract.
  • Text distractors are worse than image distractors. An equal number of irrelevant text documents hurts retrieval more than irrelevant images do — the effect the authors label Chaos in the Text.
  • The cause is a modality preference. Text representations are scored systematically higher in similarity than image representations, which is what allows irrelevant text to displace relevant images in the ranked list.
  • Trident improves mixed-modality retrieval. The abstract states gains on both CLIP-based and VLM-based architectures, evaluated across visual document and natural image benchmarks. No specific figures are given in the abstract.
  • Trident reduces fragility. The method reportedly lowers sensitivity to both modality composition and text distractors.
  • Single-modality retrieval also benefits. Average single-modality performance increases rather than being traded away for the mixed-modality gains — an important point, since a debiasing method that degrades unimodal retrieval would be less useful.

Methodology in Plain English

The authors first set up a controlled diagnostic. They take corpora and gradually swap out image documents for text that means the same thing, watching how retrieval quality moves as the mixture shifts from all-image to all-text. This isolates modality composition as a variable rather than confounding it with corpus content — the documents stay semantically equivalent, only the form changes. They then probe the ranking behavior to find where the failure comes from, which is how they arrive at the modality-preference explanation: text simply scores higher than images, so text noise wins even when the image is the right answer.

For the fix, they treat each document as having three interchangeable representations — a text view, an image view, and a fused view — and train the retriever to regard all three as valid positives for the same item rather than privileging one. The training objective, Multi-Positive View InfoNCE, couples the usual goal of separating relevant from irrelevant items with an explicit balance term so that no single view dominates the similarity space. This is a training-time intervention; it does not require changing how documents are stored or retrieved at inference.

Why This Matters

The finding reframes a problem that is easy to miss. Practitioners often evaluate retrievers on clean, single-modality benchmarks and then deploy them over messy production corpora, where the V-shaped degradation would surface as unexplained ranking failures rather than as a benchmark regression. The paper's diagnosis — that the problem is a systematic scoring bias favoring text, not merely noise — points toward a specific class of remedy rather than generic data cleaning.

Real-world applications:

  • Enterprise and legal document search, where PDFs, slides and scanned pages mix body text with figures, charts and diagrams that carry the answer.
  • E-commerce product retrieval, where listings combine product images, titles and descriptions, and text-heavy listings can crowd out visually correct matches.
  • Multimodal RAG pipelines, where a retriever selects evidence from a mixed index before a vision-language model generates an answer, so a ranking bias propagates into the final output.
  • News, social media and web search, where pages interleave images with captions and surrounding text of varying relevance.

Industry relevance: Any team building retrieval over heterogeneous content — search engines, recommendation systems, enterprise knowledge platforms, or agentic systems that fetch evidence — is affected, because the failure mode appears exactly when modalities are combined, which is the production norm rather than the exception.

Future Directions

  • Confirming the generality of the bias. The abstract covers CLIP-based and VLM-based architectures; whether the same modality preference appears in other encoder families, or in audio and video, is left open.
  • Diagnosing the origin of modality preference. The abstract establishes that text scores higher but does not say whether this comes from pretraining data statistics, the contrastive objective, or architectural asymmetry.
  • Standardizing mixed-modality evaluation. A finding that hinges on corpus composition suggests benchmark suites should report performance as a function of modality mixture, not just as single-modality scores.
  • Testing downstream consequences. Since retrieval feeds generation in RAG and vision-language pipelines, whether correcting the ranking bias improves end-task answers is a natural next question.

Target Audience

Researchers and graduate students working on dense retrieval, multimodal representation learning, and vision-language models; and practitioners — ML engineers and applied scientists — building search or RAG systems over heterogeneous document collections. Readers primarily interested in the diagnostic finding will benefit even without engaging the training method, while those implementing retrievers will find the Trident approach the actionable part.

Authors’ abstract

Dense retrievers have made significant progress on text and image corpora, but whether these capabilities extend reliably to mixed corpora containing text, image, and fused text-image documents remains unclear. In this paper, we systematically examine retrievers across architectures and find that their performance is highly sensitive to modality composition. As image documents are progressively replaced with semantically corresponding text representations, retrieval performance follows a pronounced V-shaped curve, remaining strong on single-modality corpora but degrading substantially when modalities coexist. In particular, irrelevant text causes more severe degradation than an equal number of irrelevant images, a phenomenon we term Chaos in the Text. Further analysis reveals modality preference, whereby text representations receive systematically higher similarity scores, allowing irrelevant text to outrank relevant images. To mitigate this bias, we introduce Trident, which constructs text, image, and fused text-image views of each document as co-equal positives and jointly optimizes relevance discrimination and positive-view balance through Multi-Positive View InfoNCE. Experiments across visual document and natural image benchmarks show that trident improves mixed-modality retrieval on both CLIP-based and VLM-based architectures, reduces sensitivity to modality composition and text distractors, and increases average single-modality retrieval performance.

Read the original paper