Skip to content
AI.info

Research

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Overview Research area: Text-to-Audio-Video (T2AV) generation, a multimodal generative AI task that jointly synthesizes video and audio from natural-language prompts; the paper sits at the evaluation

arXiv
2512.21094
Published
2025-12-24
Authors
Zhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang, Yuanxing Zhang, Jiahao Wang, Jialu Chen, Miao Deng, Yubin Guo, Chenxi Liao, Yize Zhang, Zhaoxiang Zhang, Jiaheng Liu

AI summary

Overview

Research area: Text-to-Audio-Video (T2AV) generation, a multimodal generative AI task that jointly synthesizes video and audio from natural-language prompts; the paper sits at the evaluation and benchmarking side of computer vision and audio research.

Technical level: Advanced. The paper assumes familiarity with generative video/audio models, embedding-based similarity metrics, and MLLM-as-a-judge evaluation protocols.

One-sentence scope: The paper introduces T2AV-Compass, a 500-prompt unified benchmark with an objective-plus-judge dual-level evaluation framework, and uses it to benchmark 15 T2AV systems.

What This Paper Is About

T2AV generation systems can now produce video and sound together from a text prompt, but there is no agreed way to measure how good they are. Existing benchmarks mostly judge video quality alone, or audio quality alone, or evaluate joint audio-video output with narrow metric sets that miss cross-modal alignment, instruction following, and perceptual realism. The authors build a single benchmark and evaluation protocol intended to test all of these at once.

Key Contributions

  1. Taxonomy-driven high-complexity benchmark. T2AV-Compass contains 500 dense prompts produced by a hybrid pipeline of taxonomy-based curation and video inversion. The prompts target fine-grained audiovisual constraints, such as off-screen sound and physical causality, that the paper says existing evaluations frequently overlook.

  2. Unified dual-level evaluation framework. The framework combines objective signal-level metrics with an MLLM-as-a-Judge protocol. Judge decisions are organized around QA checklists, which the authors present as bridging low-level fidelity and high-level semantic logic with better interpretability.

  3. Extensive benchmarking and empirical insights. The paper reports a systematic evaluation of 15 state-of-the-art T2AV systems, including proprietary models such as Veo-3.1 and Kling-2.6, and identifies an "Audio Realism Bottleneck" in which current models struggle to synthesize physically grounded audio textures that match visual fidelity.

Main Findings

  • Prompt suite scale and complexity: T2AV-Compass has 500 prompts and 13 metrics, with an average of 154 tokens, 4.03 subjects, and 3.61 events per prompt. For comparison, the paper reports VABench at 778 items, 15 metrics, and 50/3.01/2.31; JavisBench at 10,140 items, 5 metrics, and 65/3.68/1.78; VBench at 946 items and 16 metrics with 10/1.34/1.06; TTA-Bench at 2,999 items and 10 metrics with 20/2.86/1.68; Verse-Bench at 600 items and 4 metrics with 68/2.01/1.38; Harmony-Bench at 150 items and 6 metrics; and UniAVGen at 100 items and 3 metrics.

  • Prompt enrichment worked as intended: Gemini-2.5-Pro rewriting increased average prompt length from 54 to 154 tokens and the average number of constraint points from roughly 5 to 10. Human refinement reduced about 650 candidates to 400 retained rewritten prompts, and a separate video-inversion stream added 100 prompts from 100 YouTube clips of 4–10s.

  • Closed-source models lead on the compact average: Under the Average summary in the subjective results, Veo-3.1 ranks first (70.29), followed by Sora-2 (69.83), Kling-2.6 (68.16), and Wan-2.6 (67.68). The paper states this compact average should not replace dimension-level diagnosis.

  • Open-source and composed systems are competitive on parts of the task: Among open-source end-to-end models, LTX-2 is strongest overall (63.72) and achieves the best Video Realism (89.95). Wan-2.6 leads on instruction following (IF Video 78.52, IF Audio 74.95). The pipeline Wan-2.2 + Hunyuan-Foley reaches 89.63 Video Realism.

  • Audio Realism Bottleneck: The strongest model overall, Veo-3.1, has an Audio Realism score of 49.95, and the paper describes audio realism as a universal bottleneck with artifacts, muffled timbres, and weak material–timbre grounding. Audio failures are decomposed as 50.3% for biological sounds (hardest), 42.9% for mechanical sounds, and 40.5% for musical sounds (easiest); sequential prompts fail more than simultaneous ones (48.8% vs. 40.7%).

  • Dynamics is the hardest video dimension, Sound Effects the hardest audio dimension: Sub-dimension analysis shows frontier models drop notably when prompts require complex motion execution and interactions, while Speech/Music requirements are comparatively easier to satisfy than Sound Effects.

  • Failure scales with complexity: Failure rates rise from 8–10% for single subjects to 21–44% for crowds. Wan-2.6 rises from 22% failure on single-event prompts to 63% on long-narrative prompts (4+ events), while LTX-2 reaches 80% failure. The paper reports 63–80% failure on long-narrative prompts across models and describes breakdown in temporal coherence beyond roughly 4–5 events.

  • Model-specific objective results: Seedance-1.5 achieves the highest Audio Aesthetic (AA 7.403), A–V Alignment (0.2875), and Lip Sync (1.560); PixVerse-V5.5 achieves the best Temporal Synchronization (DS 0.6627) and Speech Quality (SQ 1.824); Veo-3.1 has VT 13.39; the Wan-2.2 + audio pipeline has the highest VT (13.43) but high synchronization error (DS ≈ 0.89); AudioLDM2 + MTV has the highest Text–Audio alignment (0.2698) with lower video quality scores. JavisDiT does not support speech generation.

  • Judge reliability: On a 50-prompt subset, inter-human L1 distance is 0.949 overall; Gemini-2.5-Pro is 1.087 overall and aligns most closely with humans on IF Video, IF Audio, and Video Realism, while audio realism is harder. Gemini-2.5-Flash measures 1.207 overall and Qwen3-Omni-Flash 1.473. In pairwise forced-choice comparisons across five systems (Veo-3.1, Wan-2.6, PixVerse-V5.5, Ovi-1.1, JavisDiT) fitted with a Bradley–Terry model, Gemini 2.5 Pro matches the human-derived ordering.

  • Judge stability: Repeating the full judging process three times on the same 500 video samples, Gemini 2.5 Pro has a coefficient of variation of ≤ 1.02% across subjective dimensions, and Gemini 2.5 Flash yields a five-system ranking with Spearman ρ = 0.9248.

  • Checklist and metric validation: Of 970 audio-related checklist questions, 307 (31.65%) require audiovisual synergy. Five human annotators rating AV synchronization show pairwise Spearman correlation of 0.8660, and the synchronization metric reaches SRCC = 0.9172 against mean human scores. Resampling audio from 48 kHz to 16 kHz changes T–A (CLAP), A–V (ImageBind), PQ, and SQ by only 0.0004, 0.0005, 0.0015, and 0.0015; DOVER VT at 24, 12, and 8 fps is 0.7543, 0.7519, and 0.7551.

  • Metrics are largely independent: After de-meaning z-scored metrics, absolute correlations between most metric pairs remain below 0.3, which the authors use to argue limited redundancy across objective and subjective dimensions.

Methodology in Plain English

The authors first assemble text prompts from several existing sources (VidProM, the Kling AI community, LMArena, and Shot2Story), embed them with all-mpnet-base-v2, and remove near-duplicates using a cosine-similarity threshold of 0.8, then apply square-root sampling to keep long-tail semantics. Gemini-2.5-Pro rewrites the sampled prompts to add visual, motion, acoustic, and cinematographic constraints, and humans remove non-compliant, overly long, or illogical cases. To avoid text-only hallucination, a second stream takes 100 real YouTube clips of 4–10s, generates dense temporally aligned captions with Gemini-2.5-Pro, and has humans resolve discrepancies against the source video.

Evaluation then runs on two levels. The objective level scores video quality (Video Technical score from DOVER++, Video Aesthetic score from Aesthetic Predictor V2.5), audio quality (Audio Aesthetic score defined as the arithmetic mean of Perceptual Quality and Content Usefulness, plus a Speech Quality score from NISQA), and cross-modal alignment (Text–Audio via CLAP, Text–Video via VideoCLIP-XL-V2, Audio–Video via ImageBind, Temporal Synchronization via the DeSync measure from Synchformer, and lip-sync via LatentSync). The subjective level converts each prompt into verifiable QA checklists spanning 7 primary and 17 sub-dimensions for instruction following, and adds realism checks covering motion smoothness, object integrity, temporal coherence, acoustic artifacts, and material–timbre consistency. Gemini-2.5-Pro acts as judge and must state reasoning before giving a 5-point score.

Why This Matters

Impact on research: The benchmark gives a single, taxonomy-organized testbed with 13 metrics and interpretable judge diagnostics, replacing the fragmented practice of measuring video quality or audio quality in isolation. It also supplies a concrete diagnostic claim — the Audio Realism Bottleneck — that future model work can target.

Real-world applications:

  • Automated quality assurance for systems that generate video with sound for advertising, short-form content, and virtual production.
  • Content moderation and authenticity auditing, where standardized realism and alignment scores can flag synthetic audiovisual media.
  • Accessibility and dubbing workflows that depend on speech quality and lip-sync measurements such as SQ and LS.
  • Creative tooling feedback, since per-dimension checklists tell a user which prompt constraints failed rather than only giving a single score.

Industry relevance: The evaluation covers both proprietary systems (Veo-3.1, Sora-2, Kling-2.6, Wan-2.6, Wan-2.5, Seedance-1.5, PixVerse-V5.5) and open or composed pipelines (LTX-2, Ovi-1.1, JavisDiT, and combinations of HunyuanVideo, Hunyuan-Foley, MMAudio, AudioLDM2, and MTV), so it maps directly onto the model choices available to product teams. The finding that cascaded T2V → V2A pipelines can match end-to-end models on modality-specific quality but lag in global audio-visual alignment informs build-versus-buy decisions, and the shared code and data links point to adoption as a community leaderboard.

Future Directions

  1. Extending T2AV-Compass to long-duration videos of more than 10 seconds, addressing the 63–80% failure rates observed on long-narrative prompts.

  2. Adding explicit event-level sound-role attribution, so that individual sound sources can be traced to specific visual events.

  3. Developing distilled lightweight evaluators to reduce the computational overhead of repeated MLLM-as-a-Judge evaluation during model development.

  4. Improving native audiovisual modeling with better temporal audio control and material-aware sound grounding, rather than relying only on composed pipelines.

  5. Open questions raised by the paper include reducing dependence on closed-source Gemini judge variants subject to version drift and vendor updates, and expanding prompt coverage to rare physical interactions or niche artistic concepts at the extreme long tail.

Target Audience

Researchers and engineers building or selecting T2AV generation systems, benchmark designers working on multimodal evaluation, and MLLM-as-a-judge methodologists. The paper is also relevant to applied teams that need per-dimension diagnostics (instruction following, realism, synchronization) rather than a single aggregate score, and to anyone studying physical plausibility and audio-visual grounding in generative models.

Authors’ abstract

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 11 representative T2AVsystems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.

Read the original paper