Skip to content
AI.info

Research

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Overview Research area: Long-form audio-visual question answering (AVQA), multimodal evidence retrieval, and efficient inference for omni-modal language models. Technical level: Advanced. The paper as

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
arXiv
2609.39938
Published
2026-09-30
Authors
Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng

AI summary

Overview

Research area: Long-form audio-visual question answering (AVQA), multimodal evidence retrieval, and efficient inference for omni-modal language models.

Technical level: Advanced. The paper assumes familiarity with LoRA adapters, KV-cache context limits, retrieval ranking, and audio-visual tokenization.

Scope: LEAP is a two-stage, trained evidence-retrieval framework that answers hour-scale audio-visual questions while holding per-pass context and peak memory constant in recording duration.

What This Paper Is About

Hour-long recordings create a context dilemma: encoding the whole recording exhausts the model's context limit, while uniformly compressing it dilutes the brief acoustic and visual moments that actually contain the answer. LEAP's goal is to let the answering model retrieve its own evidence, reading the timeline as a fixed grid of blocks and short candidate windows, then re-reading only the highest-scoring windows in a single bounded answer pass.

Key Contributions

  1. Evidence retrieval without a whole-recording read. LEAP narrows full-length recordings to a bounded set of evidence-rich windows, so per-pass context and working memory stay O(1) in duration. One pass over the whole recording instead costs roughly 47k tokens per hour, and the backbone's 65,536-token position limit is exhausted at about 81 minutes.

  2. The transcript as a second scanning channel. Because localization is decoupled from reasoning, the same block grid can be searched over pre-computed transcripts without decoding media frames. Only the selected windows are re-read as raw audio and video, preserving fine-grained visual and non-speech evidence a transcript misses.

  3. Historical retrieval as the stream arrives. The block grid natively supports causal queries, enabling streaming inference without streaming-specific training. The localization pass selects windows over media blocks or an online transcript index up to query time, and the answer pass re-reads them from raw media.

  4. Trained localization and answer stages. A localization LoRA improves which windows are selected, and an answer LoRA improves the answers read from the same windows; the two adapters are never stacked.

Main Findings

  • Gains over the backbone: Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5–16.8%, and transfers to MiniCPM-o 4.5, surpassing its published results by 3.1–13.0%.

  • Main benchmark accuracy (Qwen3-Omni-30B): LEAP reaches 60.9 on TraceAV, 45.8 on LVOmni, 53.7 on VideoOdyssey (1–4h), and 66.3 on MMOU, versus 56.4, 40.7, 36.9, and 56.5 for the same backbone run under the official-style recipe.

  • Ablations isolate both stages: "LEAP without retrieval," where the answer LoRA reads the whole clip in one pass, scores 54.8 / 42.9 / 41.9 / 61.8, and "LEAP without answer training," where the localization LoRA also answers the selected windows, scores 60.1 / 43.0 / 46.7 / 61.4.

  • Localization training raises evidence coverage: the share of questions whose annotated evidence the selection retains grows substantially in the answer-pass windows, and in the retained blocks most on VideoOdyssey. The trained selector also beats an untrained ranking score, an off-the-shelf retriever, or the confidence of an answer read on each block.

  • Windows must be where they are: LEAP answers significantly more accurately on TraceAV and VideoOdyssey than when the same windows are redrawn at random positions, and on VideoOdyssey it is significantly ahead of equally spaced blocks in both accuracy and evidence coverage.

  • No collapse with duration: LEAP's accuracy shows no systematic collapse as videos get longer, and on MMOU it rises with duration. It leads a context-filled montage baseline significantly in every duration quintile.

  • Transcript channel is cheaper and faster: the transcript channel spends 5.41x / 4.74x / 6.21x / 4.93x fewer localization tokens than the media channel on TraceAV, LVOmni, VideoOdyssey, and MMOU. Median query-time seconds per question are 7.6 / 11.0 / 18.6 / 3.4 for the transcript channel versus 12.4 / 15.6 / 30.0 / 5.5 for the media channel.

  • The transcript outline helps: removing the minute-stamped transcript outline from the answer pass costs LEAP accuracy significantly on TraceAV, VideoOdyssey and MMOU; its gain is largest per question where retained windows miss the evidence.

  • Causal access (StreamArena, Qwen3-Omni-30B): LEAP-media leads with 32.0 HR average versus 25.4 for whole-prefix reading and 20.2 for uniform windows, and it also leads in the farthest evidence bin (>30 minutes: 23.6 for LEAP-media, 24.7 for LEAP-transcript, 12.6 for whole prefix, 9.3 for uniform windows).

  • Audio in both passes matters: on TraceAV, silencing only the localization pass changes the selected windows substantially, and muting only the answer pass costs accuracy significantly on the hearing-required and cross-modal question classes.

Methodology in Plain English

LEAP fixes the temporal hierarchy by the clock rather than learning it. A recording is cut into non-overlapping 600-second blocks, and each block into 75-second candidate windows, giving 8 windows per full block. For each question, a lightweight localization pass reads one block at a time: video is compressed by sparse frame sampling (0.5 fps capped at 24 frames), while all 600 seconds of audio are encoded at the backbone's native rate. The block's windows are listed as lettered options in a prompt, and the model's next-token logits over those letters become window scores. A block's ranking score is the maximum window score.

The top 3 blocks are retained, and inside each, the top 3 windows; at most 9 windows enter the answer pass, restored to chronological order. Only those windows are reloaded from the source and re-encoded at a fixed per-window context — 32 frames per 75-second window plus that window's audio — together with the question, the options, absolute timestamps, and a question-ranked transcript outline capped at 4,000 tokens. The whole pipeline costs M+1 forward passes per question, independent of how many blocks are retained.

Training uses two separate rank-16 LoRA adapters (alpha = 32) on the attention query, key, value and output projections, with the backbone frozen. The localization adapter is trained by cross-entropy over option-letter logits against the candidate window that best covers an annotated evidence span, on questions derived from LongVALE; the resulting corpus keeps 40,284 of 69,630 questions over 4,735 source videos. The answer adapter is trained on 3,404 instances over 1,022 distinct gold source videos, each instance stitching one gold segment with two distractors drawn from other recordings. Checkpoints are step 2,000 for localization and step 900 for the answer adapter on Qwen3-Omni; MiniCPM-o 4.5 uses overlap-fraction BCE for localization and 13,192 answer instances over 4,215 distinct gold source videos, deploying step 1,100.

Why This Matters

Research impact: LEAP reframes long-form AVQA as a bounded-retrieval problem rather than a compression problem. It shows that localization can be trained directly within a single block — unlike selectors trained by reinforcement from the answer, which carry the whole recording and a live answering model in every update — and it demonstrates that decoupling localization from reasoning lets one block grid serve two scanning channels that share a single answer pass.

Real-world applications:

  • Searching hour-scale recordings, meetings, or broadcasts for the few minutes that answer a specific question.
  • Live or streaming question answering over an incoming feed, using a transcript index built as the stream arrives with no lookahead.
  • Audio-heavy analysis, where non-speech cues and unspoken visual detail survive because the answer pass reads raw media rather than text alone.
  • Cost-sensitive deployment, substituting the transcript channel where evidence is spoken to cut localization tokens by 4.7–6.2x and reduce latency.

Industry relevance: The method keeps deployed configuration identical across datasets and carries it unchanged to a second backbone and to causal streaming, which lowers the engineering burden of adopting it. The token accounting — roughly 47k tokens per hour against a 65,536-token position limit reached at about 81 minutes — makes the context cost of long-form omni-modal inference explicit for system designers.

Future Directions

  • Developing stronger retrieval scores within the same bounded-cost structure.
  • An answer pass that can detect when its retained windows miss the evidence and return to the scan for more.
  • Extending the transcript and media channels to settings where the evidence is not predominantly spoken, since the transcript channel's relative retention falls as less evidence is spoken.
  • Broadening evaluation of the trained selector across more backbones and benchmarks, including the video-only benchmarks reported only in the appendix.

Target Audience

Researchers and engineers working on long-context multimodal models, audio-visual question answering, retrieval-augmented generation over media, and streaming inference. It is most useful to readers who already understand LoRA fine-tuning, omni-modal tokenization, and context-limit trade-offs, and who want a concrete recipe for making hour-scale perception cost independent of recording length.

Authors’ abstract

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.

Read the original paper