Research
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
Overview Research area: Computer Vision / generative video evaluation, specifically benchmarks for text-to-video (T2V) generation. Technical level: Intermediate — the paper assumes familiarity with di
- arXiv
- 2510.26412
- Published
- 2025-10-30
- Authors
- Xiangqing Zheng, Chengyue Wu, Kehai Chen, Min Zhang
AI summary
Overview
- Research area: Computer Vision / generative video evaluation, specifically benchmarks for text-to-video (T2V) generation.
- Technical level: Intermediate — the paper assumes familiarity with diffusion and autoregressive video models, CLIP-style embeddings, and VQA-based scoring, though the framework itself is described step by step.
- Scope (one sentence): The paper introduces LoCoT2V-Bench, a 234-video benchmark with long, multi-scene prompts (average 248.85 words) and LoCoT2V-Eval, a five-dimension evaluation framework used to score 17 long-video generation methods.
What This Paper Is About
Text-to-video models now produce short clips well, but they struggle on long videos (defined here as longer than 10 seconds) with multiple scene transitions and complex dynamics. Existing benchmarks mostly use short, single-scene prompts and emphasize visual quality, temporal consistency, and overall prompt adherence, leaving fine-grained prompt alignment and higher-level thematic expression largely unmeasured. The authors build a benchmark from real-world videos with hierarchical metadata (character settings, object dynamics, camera behaviors) and an evaluation suite that scores models on perception, coarse-to-fine alignment, temporal and dynamic quality, and how well outputs fulfill human expectations.
Key Contributions
- LoCoT2V-Bench: a prompt suite built from 234 curated longer videos spanning 18 themes, featuring multi-scene prompts with hierarchical metadata. The prompts average 248.85 words and a complexity score of 8.70, versus 7.64 words / 2.54 complexity for VBench-Long and 125.46 words / 8.13 for VBench 2.0.
- LoCoT2V-Eval: a multi-dimensional evaluation suite covering perceptual quality, coarse-to-fine text-video alignment (overall and fine-grained), temporal quality, dynamic quality, and a new Human Expectation Realization Degree (HERD) metric for higher-level narrative and emotional attributes.
- Benchmarking and analysis: evaluation of 17 representative long-video generation (LVG) methods, revealing weaknesses in fine-grained text-video alignment and long-term character consistency, plus analyses showing the suite aligns with human preferences.
- Open release: code and data are released at the project's GitHub repository.
Main Findings
- Strong perception, weak fine-grained alignment: Models score well on frame-level Perceptual Quality (LongSANA reaching 84.11%) and on environment stability, with Background Consistency (BC) and Warping Error (WE) scores consistently exceeding 90% across most baselines. However, Overall Alignment (OA) to Fine-Grained Alignment (FGA) drops sharply, exposing limits in handling complex prompts.
- Character identity is the hardest temporal problem: Most methods score below 50% on Character Consistency (CC), in sharp contrast to the high BC scores, showing that maintaining subject identity is harder than keeping static backgrounds stable.
- Proprietary models lead on alignment and expectations: Kling 3.0 and Sora2 show the best overall performance, leading on HERD (up to 87.47%) and text-video alignment (Kling 3.0: OA 73.08, FGA 56.94; Sora2: OA 69.64, FGA 54.09). Their overall averages are 71.49 and 72.13 respectively.
- LongLive stands out on temporal quality: LongLive achieves the highest Temporal Quality score (83.66%) and the highest Character Consistency among listed methods (54.92%), with an overall average of 70.56.
- Top average scores: LongCat-Video averages 71.72, Sora2 72.13, Kling 3.0 71.49, Seedance 1.5-Pro 70.02, LongLive 70.56.
- Dimension correlations: Spearman rank correlations across all generated videos show TVA correlates more strongly with TQ and HERD, likely because all three depend on prompt-related semantics; aside from TVA, PQ correlates consistently strongly with DQ, TQ, and HERD, meaning higher visual quality tends to accompany better overall scores.
- Human alignment: On a randomly sampled 5% of generated videos scored by three experienced annotators, PQ achieved PLCC 71.39 / SRCC 70.20 / KRCC 54.67; Overall Alignment 63.38 / 67.23 / 52.68; Fine-grained Alignment 53.55 / 52.70 / 37.96; Character Consistency 47.17 / 49.90 / 39.19; Background Consistency 52.80 / 51.11 / 36.57; Dynamic Quality 53.86 / 54.81 / 41.01; and the HERD Score 54.86 / 58.44 / 44.92.
- Character removal hurts background consistency alignment: Masking character regions before computing BC yielded much worse human alignment (PLCC 18.63 / SRCC 30.98 / KRCC 21.59) than using complete frames, supporting the chosen approach.
- Better human alignment than VBench-Long on shared dimensions: On shared dimensions, LoCoT2V-Bench scored PQ 71.39 / 70.20 / 54.67, OA 63.38 / 67.23 / 52.68, CC 47.17 / 49.90 / 39.19, BC 52.80 / 51.11 / 36.57, compared with VBench-Long's PQ 55.10 / 50.69 / 36.77, OA 10.99 / 18.18 / 12.04, CC 41.09 / 32.65 / 22.26, BC 45.06 / 35.38 / 17.53.
- HERD design choices: The proposed dual-agent (Auditor-Evaluator) approach achieves higher and more stable human alignment across nearly all HERD sub-dimensions than QA-based evaluation or directly prompting Qwen3-VL-8B, with the exception of Visual Style, where the QA-based method aligns slightly better.
- Temporal complexity: Compared with StoryEval-Bench, LoCoT2V-Bench prompts achieve higher overall scores while containing more events on average, though not all prompts strictly follow StoryEval-Bench's temporal complexity definition.
- Prompt categories matter: Evaluated methods generally score substantially higher overall on "Daily Life" samples than on "Virtual" samples, and the best model varies across prompt categories.
Methodology in Plain English
The authors first gathered thousands of 30-60 second videos from YouTube using yt-dlp, guided by 18 predefined thematic keywords, then manually removed invalid samples (subtitles or watermarks, degraded visual quality, content-theme mismatch), leaving 234 videos evenly spread over 18 themes.
Prompts were built in stages. Seed 1.5-VL produced raw descriptions, refined iteratively with a self-refine loop. GPT-5 then expanded those descriptions into story-like prompts with detailed character settings, which humans refined for rationality, certainty, character completeness, and internal consistency. GPT-5 performed an extra automatic check, and humans re-inspected flagged samples. Post-processing tracked repeated character names and generated new ones when duplicates appeared, and refined potentially unsafe content such as depictions of blood or violence.
Evaluation runs across five dimensions. Perceptual Quality averages DeQA-Score frame scores using a multi-scale pyramid sampling strategy over non-overlapping temporal windows. Text-Video Alignment splits into Overall Alignment, scored by Qwen3-VL-8B on a 100-point scale, and Fine-Grained Alignment, a tree-structured VQA process: prompts are parsed into scenes, each scene passes an existence gate that can prune the whole subtree, and surviving scenes are checked facet by facet for Character, Background, and Camera Movement. Character checks use a state-aware adaptive anchoring mechanism — switching from a locate query to a grounded judge query once a character is confirmed — and action queries are gated by that anchoring status so unanchored characters cannot earn action credit. Background and camera facets use direct queries.
Temporal Quality combines Character Consistency (SAM3 traces character trajectories, Qwen3-VL-8B verifies instances, and FG-CLIP2 embeddings are compared by cosine similarity against the highest-confidence anchor), Background Consistency (FG-CLIP2 frame-to-frame similarity in a streaming implementation), and Warping Error (optical-flow alignment discrepancy mapped to [0,1] via exponential normalization e^(-ax)). Dynamic Quality is measured at frame, segment, and video levels, aggregating sub-dimensions with a pre-fitted linear model. HERD has GPT-5 generate structured expectations from the prompt, then uses a decoupled Auditor (produces a factual report without seeing expectations) and Evaluator (scores 1 to 5 per dimension using both report and video), with the six normalized dimension scores macro-averaged.
Why This Matters
This work shifts evaluation of long video generation from short, single-scene prompts and purely visual quality toward script-level, multi-scene prompts and higher-level narrative attributes. It gives the field a diagnostic instrument that separates where models succeed (perceptual quality, background stability) from where they fail (fine-grained prompt faithfulness, character identity), and it validates that instrument against human judgment and against an already widely adopted benchmark.
Real-world applications:
- Filmmaking and professional video production, where script-level instructions, choreographed camera movements, and multi-scene coherence must be honored exactly.
- Advertising and branded content, where characters, visual style, and thematic messaging need controlled, repeatable adherence.
- Story visualization and short-form content pipelines, where multi-shot narratives must maintain character identity across scene transitions.
- Model development and selection, where teams need a diagnostic benchmark to decide which LVG method fits a given prompt category or production constraint.
Industry relevance comes from the paper's framing of the gap between casual curiosity (short prompts, "random imagination") and professional workflows requiring deterministic control; the benchmark targets the latter, and the finding that the best model varies by prompt category matters for practical model routing.
One detail to flag: the abstract and body describe 17 representative LVG methods, while the Figure 1 caption refers to "13 LVG methods." The 17 methods listed across the multi-prompt and direct-input splits are FreeNoise, MEVG, FreeLong, DiTCtrl, StoryAdapter, CausVid, FIFO-Diffusion, SkyReels-V2, Vlogger, VGoT, SANA-Video, LongLive, LongSANA, LongCat-Video, Sora2, Seedance 1.5-Pro, and Kling 3.0.
Future Directions
- Improving prompt faithfulness and identity preservation, which the paper names as the key remaining challenge for long-form video generation.
- Closing the gap between Overall Alignment and Fine-Grained Alignment, so that models satisfy detailed scene- and attribute-level constraints rather than only global prompt gist.
- Raising Character Consistency, where most methods score below 50% while Background Consistency exceeds 90%.
- Refining high-level evaluation: the Visual Style dimension is the one case where the QA-based HERD variant aligned better with humans than the Auditor-Evaluator design, and the authors suggest textual auditor reports may introduce hallucinations there.
- Extending the framework's coverage and understanding category-dependent behavior, since the best model varies across prompt categories and scores skew higher on "Daily Life" than "Virtual" samples.
Target Audience
Researchers and engineers working on video generation, video generation evaluation, and multimodal benchmarks; model developers who need to select or diagnose long-video generation systems; and practitioners in filmmaking, advertising, or storytelling pipelines who require deterministic control over multi-scene, character-consistent output. Readers wanting only the headline result should know that strong perceptual quality and background stability coexist with weak fine-grained prompt alignment and character consistency across the 17 evaluated methods.
Authors’ abstract
Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present LoCoT2V-Bench, a benchmark for long video generation (LVG) featuring multi-scene prompts with hierarchical metadata (e.g., character settings and camera behaviors), constructed from collected real-world videos. We further propose LoCoT2V-Eval, a multi-dimensional framework covering perceptual quality, text-video alignment, temporal quality, dynamic quality, and Human Expectation Realization Degree (HERD), with an emphasis on aspects such as fine-grained text-video alignment and temporal character consistency. Experiments on 17 representative LVG models reveal pronounced capability disparities across evaluation dimensions, with strong perceptual quality and background consistency but markedly weaker fine-grained text-video alignment and character consistency. These findings suggest that improving prompt faithfulness and identity preservation remains a key challenge for long-form video generation. Our code and data are released at https://github.com/XqZeppelinhead0702/LoCoT2V-Bench