Research
OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
Overview Research area: Omni-modal (audio-vision-language) large language models, specifically the evaluation and post-training of audio-visual joint reasoning. The work spans benchmark construction,

- arXiv
- 2609.39490
- Published
- 2026-09-30
- Authors
- Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu, Yiwu Zhong
AI summary
Overview
Research area: Omni-modal (audio-vision-language) large language models, specifically the evaluation and post-training of audio-visual joint reasoning. The work spans benchmark construction, automated data generation, and reinforcement learning.
Technical level: Advanced. The paper assumes familiarity with group-based reinforcement learning (GRPO, GSPO, DAPO), on-policy self-distillation (RLSD, OPSD), and token-level advantage weighting, and it derives a pointwise mutual information interpretation of a log-likelihood contrast.
Scope: The paper introduces a benchmark (OmniReasoningBench), an automated data engine (OmniQA) that produces SFT and RL corpora, and a modality-aware self-distillation method (MFSD), then evaluates the resulting OmniReasoning-30B-A3B model.
What This Paper Is About
Existing omni-modal benchmarks, training data, and learning methods largely treat audio and vision independently, so it is unclear whether models actually combine the two. The authors show that several prominent audio-visual benchmarks remain partly solvable from vision alone, then build a benchmark, a data engine, and a learning method that force answers to depend on both modalities.
Key Contributions
-
OmniReasoningBench, a benchmark requiring both modalities. It contains 1,150 multiple-choice and open-ended questions across two settings: reasoning over video (750 questions, connecting observations across events) and reasoning beyond video (400 questions, applying knowledge derived from a video to a new scenario, figure, or numerical condition), each with a reference answer and an annotated evidence chain.
-
OmniQA, an evidence-grounded data engine. It segments videos into events, generates separate timestamped audio and visual descriptions, and constructs QA pairs with dependency chains linking observations to intermediate inferences. It also produces the released corpora OmniReasoning-SFT-112K (112,463 samples with synthesized thinking processes) and OmniReasoning-RL-19K (18,991 human-validated questions and evidence annotations).
-
Modality-Factored Self-Distillation (MFSD). An on-policy self-distillation method that scores each sampled response under four clue contexts (no clues, audio clues, visual clues, joint audio-visual clues) and separates joint-clue support from the non-additive cross-modality interaction, using both for bounded token-level advantage weighting.
-
OmniReasoning-30B-A3B. A model trained with the OmniQA data and MFSD, initialized from Qwen3-Omni-30B-A3B-Thinking, reaching 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench.
Main Findings
-
Existing benchmarks leak single-modality solutions. Qwen3.5-Plus, given visual-only inputs and no audio, still scores 57.85% on WorldSense, 68.76% on Daily-Omni, 44.90% on OmniVideoBench, and 60.43% on JointAVBench.
-
OmniReasoningBench questions genuinely require audio. In the paper's Figure 1 example, Qwen3.5-Plus with no audio drops to 17.2% on multiple-choice and 8.7% on open-ended questions.
-
Joint reasoning exceeds a single-modality oracle. On reasoning over video, joint inputs outperform an oracle that counts a question as solved if either the audio-only or the visual-only run is correct, with gains ranging from 9.9 to 26.7 percentage points across three models and both question formats.
-
Benchmark composition. Reasoning over video covers ten task types with 375 multiple-choice and 375 open-ended questions; reasoning beyond video spans nine task types with 250 multiple-choice and 150 open-ended questions.
-
Training data scale and coverage. The SFT and RL corpora span eight content domains, 25 production task types, and varied video durations.
-
Model improvements. OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving over the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points respectively.
-
Gains extend beyond the targeted benchmark. The model also improves on general and long-video benchmarks, including LVOmniBench, Video-MMMU, and Video-MME-v2.
-
The detailed per-subset accuracy table is not included in the provided content. The table is referenced as Table 1 with column headers for reasoning over video (MCQ, OE) and reasoning beyond video (MCQ, OE), but the numeric rows are truncated and therefore not reported here.
Methodology in Plain English
The work has three connected parts, tied together by explicit dependencies between audio and visual evidence.
Building data from evidence. The OmniQA engine splits each video into events and produces separate audio and visual descriptions with timestamps. Gemini-3.1-Pro annotates the timestamped audio-visual descriptions. Qwen3.8 then generates a question, a reference answer, and a dependency chain linking observations to intermediate inferences, plus candidate options for multiple-choice items. Verification is layered: structural checks on the dependency graph, caption references, and timestamps; a caption-conditioned solver to check the reference answer; a media verifier that checks each clue against its supporting audio or visual clip. To screen out modality shortcuts, every question is tested under four conditions — question-only, audio-only, visual-only, and joint audio-visual — and a candidate is kept only if the clues are valid, the joint condition is correct, and none of the restricted conditions is correct. A separate thinking generator then receives the timestamped captions, the verified QA pair, the evidence chain, and any additional figure, but not the source video, and expands the chain into observations, intermediate inferences, and a final answer with numeric calculations shown.
Improving credit assignment. Most reasoning-oriented post-training methods assign credit based on overall response quality, without distinguishing whether a prediction came from audio, vision, or their interaction. MFSD samples a group of responses without privileged clues, then scores the same response under four clue contexts. The likelihood gain over the no-clue context is computed for audio, visual, and joint clues. The interaction term subtracts the individual audio and visual gains from the joint gain, so it is positive only when the two clue sets together raise a token's likelihood by more than the sum of their separate effects, and it is zero when a single modality fully explains the joint gain. Interaction scores are centered within each response, then combined with the joint-clue support via a balancing coefficient, and converted into bounded advantage weights that strengthen positive advantages and reduce negative penalties without reversing signs. The final clipped policy objective follows RLSD. Detached scores and weights prevent gradients from flowing through clue conditioning, rollouts exclude privileged clues, and the three clue-conditioned scoring views require no extra rollouts or separate teacher parameters.
Why This Matters
Impact on research. The paper argues that ablating a modality is not sufficient to prove that a benchmark requires it, and demonstrates a screening procedure that retains only questions the joint condition solves and no restricted condition solves. It also reframes audio-visual RL credit assignment: a joint-clue likelihood gain alone does not indicate joint reasoning, because the gain may come from one modality. The interaction contrast gives a measurable, token-level signal for non-additive cross-modal support. The authors note this interpretation concerns the model's response to clue conditioning and does not certify that a reasoning step is correct.
Real-world applications:
- Accessible video interfaces and audio description, where systems must connect spoken content to on-screen events.
- Video question answering and search over long, mixed-modality content such as lectures, recordings, and livestreams.
- Education and instructional media, where a spoken explanation and a visual demonstration jointly establish a rule that must be applied to a new diagram.
- Compliance, monitoring, and forensic review of footage where the decisive evidence is split between what was said and what was shown.
Industry relevance. The data engine turns raw video into verified, evidence-linked supervision automatically, which matters for teams that cannot hand-annotate cross-modal reasoning traces at scale. The method adds no extra rollouts and no separate teacher model, so it fits existing group-based RL pipelines. The reported transfer to long-video and general video benchmarks suggests cross-modal evidence optimization has value beyond the audio-visual setting itself.
Future Directions
- Step-level correctness. The interaction signal measures the model's response to clue conditioning but does not certify reasoning-step correctness, leaving room for verifiers or step-level supervision on top of the token-level weights.
- Reducing dependence on proprietary generators. The pipeline uses Gemini-3.1-Pro for annotation and Qwen3.8 for QA and thinking generation; how far the engine can be reimplemented with open models is not established in the provided content.
- Scaling and broadening the data. The paper reports eight content domains and 25 production task types; extending coverage, duration ranges, and the reasoning beyond video setting to further formats is a natural next step.
- Generalizing MFSD. Whether the modality-factored decomposition transfers to other backbones and to other modality pairings is not reported in the provided content.
Target Audience
Researchers and engineers working on omni-modal or video-language models, reinforcement learning post-training, and benchmark design. It is most useful to readers already comfortable with group-based RL objectives and likelihood-based self-distillation who want a concrete recipe for forcing and measuring cross-modal dependence, and to practitioners building data pipelines for multimodal training.
Authors’ abstract
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-visual joint reasoning, together with time-stamped clue chains that guide the annotation of thinking process. Besides our benchmark, this engine produces training data OmniReasoning-SFT-112K and OmniReasoning-RL-19K. Finally, we propose an on-policy self-distillation method Modality-Factored Self-Distillation (MFSD). It evaluates each sampled response under modality-specific clue contexts, disentangling the contributions of individual clues and their cross-modal interactions for token-level credit assignment. With our training data and learning method, our model OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points, respectively. Moreover, it delivers substantial gains on general and long-video benchmarks, including Video-MME-v2. We hope our work offers a solid step for facilitating future research in omni-modal joint reasoning.