Research
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
Overview Research area: Evaluation of multimodal large language models (MLLMs) on audio-visual (omnimodal) video captioning — computer vision and multimodal benchmarking. Technical level: Advanced. Th
- arXiv
- 2610.12458
- Published
- 2026-10-08
- Authors
- Zhongyu Yang, Jiale Tao, Ruitao Chen, Zuhao Yang, Yingfang Yuan, Xueliang Zhao, Auden, Kai Wang, Shuai Shao, Biao Wang, Steve Yves, Qinglin Lu
AI summary
Overview
Research area: Evaluation of multimodal large language models (MLLMs) on audio-visual (omnimodal) video captioning — computer vision and multimodal benchmarking.
Technical level: Advanced. The paper assumes familiarity with MLLM captioning pipelines, evaluation metrics, LLM-as-judge protocols, and temporal grounding concepts such as temporal intersection-over-union.
Scope: The paper proposes OmniCapBench, a benchmark and diagnostic framework that replaces free-form caption scoring with sets of atomic, verifiable audio-visual evaluation units scored by deterministic rules plus tightly bounded LLM semantic checks.
What This Paper Is About
Current audio-visual caption benchmarks force a trade-off: whole-caption protocols score an entire generated paragraph with a global LLM judge, giving broad coverage but no way to localize errors, while probe-based protocols (QA probes, cloze tasks) localize errors but cover only sparse facts. Because model captions are unstructured monolithic text, any scoring operator must first reverse-engineer structure and facts from prose — a process the authors show is unstable and easily fooled by fluent writing. The paper's goal is to reframe audio-visual caption evaluation as a deep-structured diagnostic framework in which the caption is natively expressed as atomic, verifiable units (entity references, visual shots, audio events) so that structure can be verified by deterministic rules and semantics judged only on aligned, isolated fields.
Key Contributions
- A Deep-Structured Evaluation Paradigm. The evaluated caption is decomposed into atomic, verifiable units, explicitly isolating semantic content, temporal grounding, identity tracking, and cross-modal association, rather than imposing a superficial format on global text.
- The OmniCapBench Benchmark. 786 densely annotated videos yielding 5,818 entities, 6,537 audio events, and 11,419 visual shots across native Reference, Shot, and Event tracks (totalling 12.8 hours, 20 level-1 and 125 level-2 categories, plus 39,160 subshots and 5,370 dialogue events).
- A Decoupled Two-Stage Scoring Pipeline. Deterministic rules verify structural integrity (valid IDs, temporal bounds, cross-links), after which bounded, localized LLMs assess semantic equivalence only on structurally aligned unit pairs, with bidirectional Recall (penalizing omissions) and Precision (penalizing hallucinations) scoring.
- An Empirical Audit of Holistic Scoring. Experiments across models, metrics, and generation formats show that global text-level scores mask localized errors and reward structurally flawed outputs, and that unconstrained LLM judges introduce severe ranking instability.
Main Findings
- Schema adherence is essentially solved. Schema Generation Compliance (SGC) is high across frontier models, reaching 97.36 for Gemini 3.1-Pro and 96.73 for Gemini 2.5-Pro, and format errors account for only 1.3% of Gemini 3.1-Pro's failure distribution. Open-source MiniCPM-o-2.6 is a clear outlier at 72.13.
- Basic visual perception is strong, but identity tracking across cuts breaks down. Gemini 2.5-Pro reaches Ref Subject F1 of 80.70 and Ref Scene F1 of 83.65, and the paper reports 84.80% Ref Subject F1 on videos under 1 minute (Appendix Table B.5). Cross-Shot Coreference Consistency (CCC), however, drops to 37.81 for Gemini 2.5-Pro and collapses below 12% for open-source models.
- Audio-visual binding is the weakest frontier capability. Gemini 3.1-Pro falls from 63.32 Event F1 to 51.46 Event-Shot Association (EVSA) F1, and open-source models fail to surpass 20% EVSA F1. Seed2.0 emits almost no audio-event units (4.54 Event F1, 3.52 EVSA F1) despite top-ranked speaker attribution at 91.80 Speaker F1.
- Speech is easy, environmental sound is not. Dialogue semantics sit at or above 80% across models, while non-dialogue audio precision is far lower (for example, 41.49 for Gemini 3.1-Pro, 4.84 for MiniCPM-o-2.6).
- Localized scoring exposes a precision-recall trade-off hidden by holistic scores. The paper reports Gemini 3.1-Pro as adopting a conservative strategy and MiniCPM-o-2.6 as an extreme case of high precision paired with low recall. Note that the narrative text and Table 4 attach these figures to different fields: Table 4 lists Gemini 3.1-Pro Subject Recall 59.05 / Precision 43.78 and Scene Recall 67.07 / Precision 51.78, and MiniCPM-o-2.6 Scene Recall 63.87 / Precision 20.39 with Subshot Precision 76.32 / Recall 27.62, whereas the prose cites 67.07% precision with 43.78% recall on subjects and 76.32% scene precision with 20.39% recall.
- Unconstrained LLM judges are unstable. Scoring identical predictions with weak (Qwen3.6-27B), mid (GPT-4o), and strong (Gemini-2.5-Pro) judges inflates Holistic-style scores from 71.24 to 93.92 and moves QA-style scores from 61.17 down to 56.57. OmniCapBench stays nearly flat across the same judges: 56.82, 58.24, 57.60.
- Global metrics mask structural errors. On hard-to-distinguish subsets of 200 videos drawn from UGC-VideoCap and video-SALMONN 2, where weaker and stronger models appear near-identical (~50% vs. ~50%) under text-centric metrics, structural constraints separate them: stronger models rise to 74.2% (aligning with human audit) while weaker models collapse to 25.8%.
- Failure signatures differ between model classes. For Gemini 3.1-Pro, failures concentrate in action hallucination (28.5%), A-V misalignment (23.7%), and audio hallucination (17.9%), whereas open-source models degrade broadly and uniformly across basic perception and complex audio-visual binding.
- Specialized captioning models were excluded. ASID-Caption and AVoCaDO are excluded from the main evaluation because supervised fine-tuning toward free-form text leads to poor instruction following on structured generation tasks (Appendix B.5).
Methodology in Plain English
The authors first choose a structured representation of video content, building on the Multi-Stream Scene Script (MTSS) design but refining it by splitting visual shots into finer subshots and replacing discrete point timestamps with continuous time ranges. Each video is stored as a reference system with three interdependent tracks: References (persistent identities for people, objects, and scenes), Events (auditory events and dialogue with precise time boundaries, e.g. a dog bark from 10.5s to 12.0s), and Shots (visual segments that serve as the unifying structure by linking references and events to specific frames).
Construction runs in three stages. Stage I curates candidate videos from open datasets and public platforms, filtering on audio presence, duration, quality, and bucket-specific density, using duration buckets of under 1 minute, 1–3 minutes, and 3–5 minutes. Stage II builds units in dependency order (References, then Events, then Shots) through an iterative Local Refinement loop: generate, evaluate and validate, locally refine, select. Programmatic validators enforce schema constraints and trigger regeneration or discarding, and textual refinement is only permitted while identifiers, timestamps, and cross-links stay strictly unchanged. Stage III applies programmatic and targeted human audits to finalize the reference system. Red timestamps burned into frames are used solely as an annotation aid; evaluated models always receive original frames.
Evaluation is split into three tasks: predicting native structured units as JSON, verifying structure and matching predictions to ground truth (shots and events by temporal IoU, references and subshots by a bounded LLM-assisted matcher restricted to local candidate pairs), and localized semantic scoring where an LLM compares only isolated fields. For open-vocabulary fields, matching runs bidirectionally — ground truth to prediction (Recall, penalizing omissions) and prediction to ground truth (Precision, penalizing hallucinations) — with F1 as the harmonic mean. Reported metrics cover Visual (RefUse, CCC, Ref Subject/Scene F1, Shot/Subshot F1 and tIoU), Audio (Event F1 and tIoU), and Audio-Visual (Speaker F1, EVSA F1) dimensions.
Why This Matters
Impact on research. The paper argues that monolithic text scoring, not model capability alone, limits diagnostic resolution in multimodal evaluation. By defining units that are natively verifiable and by restricting LLMs to bounded semantic comparisons, it offers a scoring contract that is stable across judges and traceable to specific error types, which could shift benchmark design from single scalar leaderboards toward structured diagnostic reports.
Real-world applications:
- Video search and indexing systems that need to attach sounds to the right on-screen source, rather than transcribing audio and detecting visuals independently.
- Accessibility and captioning tools, where hallucinated attributes (such as unsupported "protective eyewear" in the paper's case study) create real misinformation risk for viewers.
- Video editing and media asset management pipelines that depend on accurate shot boundaries and event timestamps for cutting and retrieval.
- Model development and QA for omnimodal assistants, where per-capability failure signatures give teams a prioritized list of what to fix.
Industry relevance. The benchmark is released by Hunyuan, Tencent, with co-authors from Nanyang Technological University and Northumbria University, and evaluates commercial and open systems including Gemini, Qwen, Seed, MiMo, and MiniCPM families. Its finding that audio and visual streams are still processed as unaligned pathways is a concrete target for product teams building omnimodal assistants, and the instability of LLM-as-judge scores is directly relevant to anyone relying on judge-based automated evaluation.
Future Directions
- Closing the identity-tracking gap. CCC collapsing to 37.81 for Gemini 2.5-Pro and below 12% for open-source models points to long-term visual object permanence as an unresolved architectural problem.
- Unifying audio and visual pathways. Failing to surpass 20% EVSA F1 for open-source models, and Gemini 3.1-Pro dropping from 63.32 Event F1 to 51.46 EVSA F1, raises the question of how to train cross-modal binding rather than parallel perception.
- Reducing deep hallucinations. Action hallucination (28.5%) and A-V misalignment (23.7%) dominating Gemini 3.1-Pro's failures suggest targeted evaluation and training for fine-grained action details and correct sound-to-shot assignment.
- Adapting specialized captioning models. ASID-Caption and AVoCaDO were excluded for poor instruction following on structured generation, leaving open how free-form-tuned captioners could be adapted to native atomic output.
- Boundaries of the evaluation scope. The paper's appendix contains a Limitations and Broader Impact section (Appendix C), but the truncated content provided does not report its specific findings.
Target Audience
Researchers and engineers working on omnimodal or multimodal LLMs, audio-visual captioning, and video understanding who need fine-grained diagnostics rather than aggregate scores. It is also relevant to benchmark designers and evaluation practitioners interested in replacing unconstrained LLM-as-judge scoring with deterministic verification plus bounded semantic comparison, and to product teams building video search, captioning, or accessibility systems where hallucinated or temporally unsupported content carries real cost.
Authors’ abstract
Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio--visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.