Research
VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
Overview Research area: Computer vision and generative AI, specifically evaluation of video generation models with a focus on their ability to render readable, correct text inside generated scenes. Te

- arXiv
- 2610.01499
- Published
- 2026-10-01
- Authors
- Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu
AI summary
Overview
Research area: Computer vision and generative AI, specifically evaluation of video generation models with a focus on their ability to render readable, correct text inside generated scenes.
Technical level: Intermediate. The benchmark design and failure analysis are accessible to a general technical reader, though the evaluation protocol relies on word error rate, chain-of-query scoring, and vision–language model judgments.
Scope: The paper introduces VTR-Bench, a 300-prompt benchmark spanning five application scenarios that separately measures text fidelity and scene/motion adherence in generated videos, evaluates 11 state-of-the-art models, and proposes a keyframe-guided agentic framework that reduces text errors.
What This Paper Is About
Video generation models can now produce realistic-looking video from text prompts, but existing benchmarks mostly judge visual quality, aesthetics, and physical plausibility while ignoring whether the text visible inside a scene is actually correct. A video can look convincing and satisfy every scene requirement yet show misspelled words, repeated text, or illegible passages, which breaks its usefulness in settings like advertisements, scientific demonstrations, and interfaces. VTR-Bench was built to measure this overlooked capability directly, by embedding prescribed text in concrete scenarios and scoring both the text and the surrounding scene and motion requirements.
Key Contributions
-
A scenario-grounded benchmark for visual text rendering. VTR-Bench places prescribed text in concrete scenes across five application scenarios — advertising, science, user interfaces, culture, and daily life — with 300 prompts (60 per category), 1,202 annotated text blocks, and explicit requirements for both the text content and its carrier (the physical or interface surface where the text should appear).
-
A decoupled, automated evaluation pipeline with human alignment checks. Text fidelity is measured by transcribing text from each specified carrier with a vision–language model and comparing it against reference strings using word error rate (WER, with α = 5 as the default setting), while scene and motion requirements are measured through a prompt-specific chain of query (CoQ) of 20 questions producing a Video Score.
-
A Keyframe-Guided Agentic Framework. A Director agent coordinates image generation, motion planning, and video generation, using visual feedback to refine image prompts, edit or regenerate first frames, and select candidates before video generation.
-
A systematic evaluation and failure analysis. Experiments on 11 state-of-the-art models quantify text rendering errors, contrast text fidelity against Video Score, analyze the effects of generation resolution and duration, and categorize transcription outcomes across 13,213 text targets.
Main Findings
-
Text rendering is broadly unsolved: the best model reaches an overall WER of 0.250. Among open-source models, Minimax H3 stands out with an overall WER of 0.447 and a Video Score of 0.756, while Wan2.2-5B, LTX-2.3, Lingbot-Video-Dense, Lingbot-Video-MOE, and HunyuanVideo-1.5 all record overall WERs between 0.896 and 0.996. Among proprietary models, Wan3.0 leads every scenario (overall WER 0.250, Video Score 0.849), followed by Seedance2.5 (0.641), while ViduQ3 (0.950), HappyHorse1.1 (0.880), and Kling v3.0 (0.979) still produce substantial text errors.
-
High Video Scores do not imply accurate text. Kling v3.0 and Seedance2.5 both score approximately 0.79 on video requirements, yet their WERs are 0.979 and 0.641 respectively. Minimax H3 renders text more accurately than most proprietary models despite a lower Video Score (0.756).
-
The motivating example shows the gap concretely. An approximately 10-second video generated by Seedance2.5 achieves a Video Score of 0.90 on the checklist of scene and motion requirements while its WER reaches 0.552, with misspelled words, repeated and incorrect text, and largely illegible passages.
-
Higher resolution helps only when the model already has some text-rendering ability. For Minimax H3, higher resolution reduces WER at both durations tested (high resolution 1344×768 vs. low 832×480). LTX-2.3 maintains a WER close to 1 across all four resolution-and-duration configurations (high 1280×704 vs. low 768×512; long 10 seconds vs. short 5 seconds). Shortening duration provides no consistent improvement in text fidelity for either model.
-
Unreadable text is more common than missing carriers. Across 13,213 text targets, the evaluator reports unreadable text for 23.31% of targets versus 9.34% for carriers not found. Wan2.2-5B and LTX-2.3 exceed half of targets judged unreadable, consistent with their high WERs.
-
Transcribability does not establish accuracy. Kling v3.0 has a transcription rate of 79.62% but an overall WER of 0.979, whereas Minimax H3 combines a transcription rate of 85.61% with a WER of 0.447.
-
Test-time refinement improves both metrics. With Minimax H3, the agentic framework reduces overall WER from 0.4468 to 0.3015 (a 32.5% relative reduction) and raises Video Score from 0.7562 to 0.8315, improving both metrics in all five scenarios relative to direct generation. The full framework reduces overall WER by a further 9.4% relative to first-frame conditioning (I2V) and achieves the highest Video Score in every scenario, with its WER advantage over I2V largest in scientific videos; in UI scenes, I2V achieves the lowest WER while the full framework achieves the highest Video Score.
-
Automated evaluation agrees closely with humans. On 100 sampled videos covering 2,000 CoQ judgments and 404 text blocks, the evaluator matches the human majority vote on 92.15% of queries. For text-block WER, Pearson is 0.9541, Spearman 0.9370, Lin's CCC 0.9536, and MAE 0.0545; at the video level, Pearson is 0.9842, Spearman 0.9731, Lin's CCC 0.9835, and MAE 0.0407.
Methodology in Plain English
The benchmark was built in stages. Human experts first defined five high-level scenario categories, then GPT-5.6 generated diverse scene seeds describing the core setting, the purpose of the text, and candidate carriers. Human reviewers screened and deduplicated those seeds, and DeepSeek-V4-Flash turned them into complete prompts specifying the exact visible strings, where each string appears, and how the text relates to the activity in the scene. To check that each prompt was visually feasible, three reference images generated by Qwen-Image were reviewed jointly by Qwen3.8-27B for carrier coverage, text hierarchy, and integration with the activity; recurring structural problems attributable to the prompt triggered revision, and three consecutive unsuccessful checks triggered reconstruction of the candidate. Validated prompts then passed a final human review. The result is 300 prompts with paired reference strings and carrier descriptions.
Evaluation keeps two questions separate. To measure text fidelity, a vision–language model transcribes each specified carrier from its clearest appearance in the video, preserving rendering errors and leaving unreadable or missing targets as empty strings; transcriptions and reference strings are tokenized with a shared deterministic tokenizer that preserves case, content-bearing symbols, and individual CJK characters while ignoring ordinary prose punctuation, and compared with a bounded token-level word error rate. To measure scene and motion fulfilment, each prompt is paired with a 20-question chain of query generated by GPT-5.6 and reviewed by human annotators, spanning Scene Attributes, Motion Adherence, Spatial Relationship, Entity Presence, and Temporal Consistency; each question is answerable yes or no, and the Video Score is the proportion of satisfied requirements.
For the agentic framework, a Director agent builds an image prompt, requests candidate first frames, and uses a vision–language model to inspect them for text accuracy, legibility, carrier coverage, and scene consistency. It can compare candidates, refine the prompt, or edit a candidate with a targeted instruction. The chosen first frame, the original prompt, and a motion plan (subject actions, camera movement, temporal progression) condition video generation using the same model weights as direct generation. A vision–language model then reviews sampled frames for text stability, carrier persistence, motion coherence, and content adherence, guiding motion-plan refinement and final video selection. In the refinement experiment, Qwen3.7-plus served as both Director agent and vision–language model, and Qwen-Image-3.0 as the image generator, with the first frame of each generated video removed before evaluation for both I2V and Agentic settings.
Why This Matters
Impact on research: The paper argues that text is an under-measured dimension of video generation quality and that a single high Video Score can hide incorrect written content. By separating text fidelity from scene and motion compliance, it provides a metric design and a failure taxonomy (carrier not found, text unreadable, transcribed-but-wrong) that other benchmark builders and model developers can reuse. The automated pipeline, validated against human annotations, also offers an alternative to purely human evaluation of on-screen text.
Real-world applications:
- Advertising, where fictional identities and product copy must appear correctly in generated promotional video.
- Scientific videos, where explanations, derivations, numerical values, and evidence must be reproduced faithfully.
- User interfaces, where credible task-specific interface text and labels convey instructions.
- Culture and daily life footage, where writing appears within practices, performances, transmission, authorship, and personal or household activities — for example product introductions, numerical values, and instructions generally.
Industry relevance: Video generation is increasingly deployed in commercial content pipelines where a misspelled product name or a wrong number changes the message. The paper's finding that the strongest evaluated model still records an overall WER of 0.250 indicates that text rendering remains a practical blocker, and its demonstration that inference-time refinement reduces WER by 32.5% with the same video generator suggests a deployment-side mitigation that does not require retraining.
Future Directions
- Improving underlying text-rendering capability rather than only configuration. The resolution and duration study shows that better settings only help models that already have some ability to render text, leaving open how models such as LTX-2.3 could be strengthened at their core.
- Understanding how scene complexity, textual density, and motion affect fidelity. The conclusion frames explicit content requirements as evaluation targets and identifies these factors as a basis for further study; the benchmark spans 58 to 496 tokens per video and up to six text blocks per prompt, providing material for such analysis.
- Extending agentic refinement across models and settings. The test-time refinement experiment was conducted with Minimax H3 only, with a Director agent and evaluator based on Qwen models, so whether the gains generalize to other generators and agent choices is untested.
- Scaling and broadening automated evaluation. The authors note that prior text-focused benchmarks rely partly on human evaluation, making repeated assessment labor-intensive; VTR-Bench reduces this dependence but still uses human review in dataset construction and alignment, and the alternative-evaluator results with Qwen3.6-27B raise the question of how sensitive model rankings are to the choice of evaluator.
Target Audience
This paper is most useful for video generation researchers and benchmark designers, model developers working on text rendering or controllable generation, and practitioners deploying generated video in text-critical contexts such as advertising, education, and interface mockups. It is also relevant to researchers studying vision–language model evaluation, since it reports detailed human-alignment statistics for automated transcription and question-answering based scoring.
Authors’ abstract
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.