Skip to content
AI.info

Research

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

Overview Research area: Information retrieval and retrieval-augmented generation (cs.IR), specifically post-retrieval evidence compression for multimodal RAG over heterogeneous text, table, image, and

arXiv
2610.00923
Published
2026-10-01
Authors
Hyojeong Yun, Jueun Kim, Wook-Shin Han

AI summary

Overview

Research area: Information retrieval and retrieval-augmented generation (cs.IR), specifically post-retrieval evidence compression for multimodal RAG over heterogeneous text, table, image, and video corpora.

Technical level: Advanced. The paper assumes familiarity with RAG pipelines, embedding-based retrieval, hierarchical representations, pairwise ranking losses, and LoRA fine-tuning.

Scope (one sentence): The paper introduces Canopy, a framework that compresses retrieved multimodal items by selecting regions at adaptive granularities within each item and pairs that compression with critic-guided additional retrieval, evaluated on five QA benchmarks over a 33M-item corpus.

What This Paper Is About

Multimodal RAG systems decide which items to retrieve, but not how much of each item the reader should see. Passing whole items bloats the context window and adds distracting content, while uniformly fine-grained selection can strip away context needed to interpret the evidence, and different regions within the same item may need different amounts of surrounding context. Canopy addresses this by representing each retrieved item as a hierarchy of original regions and using a fine-tuned node encoder plus a shared parent-relative comparison rule to decide, region by region, how deep to descend, with a critic requesting targeted follow-up retrieval when the accumulated evidence is judged insufficient.

Key Contributions

  1. Problem formulation: The authors frame post-retrieval evidence compression for heterogeneous multimodal RAG, focusing on adapting the retained evidence granularity within each retrieved item rather than fixing one granularity per item.
  2. The Canopy method: A framework combining a node encoder fine-tuned on gold evidence with parent-relative hierarchical refinement, which selects multiple regions at different granularities and requires no LLM calls for node-level pruning.
  3. An end-to-end evaluation: Five QA benchmarks (NQ, HotpotQA, OTT-QA, MMQA, LVBench) over a heterogeneous corpus of approximately 33M text, table, image, and video items, including initial and follow-up retrieval rather than assuming relevant evidence is given.
  4. Component ablations: Ablations that separate the contributions of additional retrieval and compression to answer quality and reader-input evidence volume, including a direct measurement of encoder fine-tuning's effect on refinement decisions and pruning rates.

Main Findings

  • Higher average accuracy: Averaged over the five benchmarks (mean of EM, accuracy for LVBench), Canopy reaches 39.4 with Qwen3-VL-8B-Instruct and 37.8 with InternVL3.5-8B, versus 32.9/33.0 for Vanilla RAG, 32.7/28.9 for IRCoT, and 32.1/30.7 for UniversalRAG.
  • Gains concentrate on multi-hop QA: With Qwen3-VL-8B-Instruct, Canopy is best in EM and F1 on the four non-video benchmarks, but its +0.3 F1 over IRCoT on HotpotQA is not significant (95% CI [-1.7, +2.2]). The larger gaps appear on HotpotQA and OTT-QA, where single-round retrieval misses evidence identifiable only after earlier evidence is read.
  • Token savings with comparable accuracy: In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2–27.7% relative to the same iterative pipeline without compression, with EM changes from -0.9 to +1.2 points and a +0.4-point LVBench accuracy change; every one of those 95% CIs includes zero.
  • Additional retrieval drives accuracy: Removing iterative retrieval lowers HotpotQA and OTT-QA EM by 10.1 and 14.9 points respectively, so the multi-hop gains come from additional retrieval rather than compression.
  • Routing is complementary, not universally better: Routing helps only on LVBench, where text queries embed far from video items; Canopy with routing reaches the best LVBench accuracy (43.4 on Qwen3-VL-8B-Instruct) while unrouted Canopy stays ahead elsewhere. With InternVL3.5-8B on LVBench, routed Canopy is lower than UniversalRAG (34.0 versus 37.2) while using more evidence tokens (31,218 versus 27,290).
  • Fine-tuning sharply increases pruning: Fine-tuning raises the share of correct parent–child refinement decisions from 28.3–43.0% to 34.3–71.9% across benchmarks and raises the share of pruned items from 30.4–39.6% to 85.2–95.6%, reducing reader-input tokens by 6.8–19.1% while F1/accuracy changes by +0.3 to +0.9 points (all 95% CIs include zero).
  • Selection strategy comparison: Given gold items, Canopy exceeds both leaf-level and any-level flat selection on all five benchmarks while retaining more tokens; against keeping the whole gold item it reduces LVBench tokens by 40.0% at comparable accuracy (56.4 versus 56.2, 95% CI [-3.4, +3.8]) and saves 27.5% on OTT-QA with F1 falling from 70.5 to 67.8, while NQ and HotpotQA savings are approximately 5%. Gold Span is not an upper bound: its LVBench accuracy is 53.0.
  • Efficiency versus competing compressors: LongLLMLingua saturates 1.4–16.5 F1 below Canopy, and AKS spends 1.5x its tokens for 2.8 fewer accuracy points. S2G-RAG reaches comparable F1 with far fewer tokens but makes 5.6 LLM calls per question against 2.8 for Canopy in that comparison.
  • Evidence volume within the pipeline: At 2.1 rounds on average, Canopy passes at most 28% more evidence tokens than Vanilla RAG and 46–88% fewer than UniversalRAG.
  • Sensitivity: A larger branching factor does not consistently improve quality or reduce tokens, and retrieving more items generally increases tokens with modest or non-monotonic quality changes.

Methodology in Plain English

Canopy starts after retrieval. Each retrieved item is turned into a tree whose nodes are original regions, not generated summaries: for text the leaves are sentences, tables are partitioned by row with column headers preserved at every level, videos are split along the temporal axis into 30-second leaf segments with frames sampled at one frame every 10 seconds up to 32 frames, and images remain single nodes with no spatial pruning. Because aggregate questions over a table may need an arbitrary subset of rows, an LLM classifies the original question once, from the question alone, as aggregate or lookup; aggregate questions keep the table whole and lookup questions use hierarchical refinement.

A frozen query encoder and a fine-tuned node encoder produce query–region similarity scores. Refinement begins at the root: at each internal node, if any child scores at least as high as its parent, the parent is replaced by all qualifying children and the remaining child subtrees are discarded; if no child qualifies, the parent is kept whole. A visited leaf is retained directly. Because branches stop independently, the selected nodes form an evidence forest that can contain regions at different granularities, and no retained-unit count or token budget is prescribed. Node embeddings do not depend on the query, so they can be computed once per item and cached; node-level pruning needs only similarity comparisons, not generative LLM calls.

The node encoder is trained with a pairwise ranking loss over preferences derived from gold evidence annotations. For each internal node that an ideal refinement would visit, the method prefers gold-overlapping children over the parent, the parent over non-gold children, and gold-overlapping children over their non-gold siblings. Only LoRA parameters in the language backbone are updated; the query encoder stays frozen.

Compression alone cannot supply evidence that was never retrieved, so an LLM critic inspects the accumulated evidence against the original question. If it judges the evidence sufficient, the reader answers; otherwise it identifies a missing fact and issues a targeted follow-up query, and newly retrieved items are compressed before being added (re-retrieved items are skipped). The loop ends when the critic is satisfied or the maximum number of retrieval rounds is reached. Defaults are branching factor B = 2, k = 10 items retrieved per round, and at most R = 3 rounds, with Qwen3-VL-Embedding-2B shared as the retrieval encoder across all retrieval-based methods and a single LoRA adapter shared across all five benchmarks.

Why This Matters

Impact on research: The paper separates two decisions that multimodal RAG work often conflates: which items to retrieve and how much of each item to keep. By showing that a single shared refinement procedure transfers across text, tables, and video while needing no LLM calls for pruning, it offers a modality-agnostic alternative to modality-specific compressors and a complement to corpus selection and query routing. The ablation also gives a clean attribution: multi-hop accuracy gains trace to additional retrieval, while compression controls evidence volume.

Real-world applications:

  • Enterprise assistants that answer questions over mixed document collections containing prose, financial tables, and scanned images, where sending whole documents to a model is impractical.
  • Long-video question answering over lecture, meeting, or surveillance archives, where only a segment of a multi-hour video is relevant.
  • Open-domain question answering services that must ground responses in external knowledge newer than or absent from model training data.
  • Agentic pipelines with tight context or cost budgets, where pruning by embedding similarity rather than LLM calls reduces per-query inference cost.

Industry relevance: Retrieval pipelines pay for context twice, in latency and in token cost, and the paper reports both: #Tok and latency per question under batched execution. The finding that node-level pruning needs only similarity comparisons means the compression step is cheap to deploy alongside an existing retriever, while the critic-guided loop adds LLM calls (3.0 per question for unrouted Canopy versus 1.0 for Vanilla RAG) that practitioners must weigh against quality.

Future Directions

  • Adaptive branching and retrieval sizing: The sensitivity study finds no consistent benefit from larger branching factors and only modest or non-monotonic quality changes from retrieving more items, leaving open how B and k should be chosen per query or modality.
  • Routing versus no routing: Routing helps only on LVBench and hurts elsewhere in these experiments, including routed Canopy trailing UniversalRAG on InternVL3.5-8B on LVBench, so when to route across modality-specific corpora remains unresolved.
  • Evidence sufficiency guarantees: The critic's judgment is described as an operational stopping criterion rather than a guarantee of completeness, and reaching the round limit does not establish sufficiency, so more reliable sufficiency estimation is an open question.
  • Generalization beyond the evaluated settings: The paper reports in-domain results across five benchmarks and describes out-of-domain benchmarks (2WikiMultiHopQA, TAT-QA, and others in Appendix K), along with a routed-LVBench result where compression saves 33.3% of tokens but lowers accuracy by 1.8 points; broader transfer and the conditions under which compression preserves answer quality are not fully settled.

Target Audience

Researchers and engineers working on retrieval-augmented generation, multimodal retrieval, or context compression will benefit most, particularly those building pipelines over heterogeneous corpora who need within-item granularity control rather than another retrieval granularity choice. Readers should be comfortable with embedding similarity, hierarchical indexing, ranking losses, and LoRA fine-tuning to follow the method and ablation details.

Authors’ abstract

Multimodal RAG retrieves text, tables, images, and videos, but choosing a retrieval granularity does not determine how much context to retain within each item. Coarse units include irrelevant content, while uniformly fine selection can remove context needed to interpret the evidence. Existing compressors address this trade-off with modality-specific mechanisms, leaving open a shared procedure for adapting the retained extent region by region across heterogeneous items. We introduce CANOPY (Canonical Projection over Hierarchy), a framework for adaptive-granularity post-retrieval evidence compression. CANOPY represents retrieved items as hierarchies and uses a node encoder fine-tuned on gold evidence to score regions against the query. Parent-relative refinement compares these scores to select multiple regions at different granularities without LLM calls for node-level pruning. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added. Across five QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. Ablations indicate that additional retrieval drives the main accuracy gains on multi-hop QA. In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2-27.7% relative to the same iterative pipeline without compression, with comparable answer accuracy.

Read the original paper