Skip to content
AI.info

Research

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

Overview Research area: Computer vision, specifically text-to-image (T2I) generation and evaluation of visual text rendering. Technical level: Intermediate. The paper is a benchmark and evaluation stu

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
arXiv
2610.09823
Published
2026-10-07
Authors
Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li, Jun Xie, Jiacheng Liu, Jungang Li, Yu Huang, Xuanyi Liu, Yue Ding, Zecheng Wang, Lei Zhao, Mingda Wang, Zhenglin Cheng, Peng Sun, Tao Lin

AI summary

Overview

  • Research area: Computer vision, specifically text-to-image (T2I) generation and evaluation of visual text rendering.
  • Technical level: Intermediate. The paper is a benchmark and evaluation study; understanding it benefits from familiarity with diffusion/flow image generators, OCR metrics, and vision-language model (VLM) judging, but the core argument is stated in plain terms.
  • Scope in one sentence: The paper introduces UltraText Bench, a bilingual (English/Chinese) benchmark of 432 prompts across 24 scene categories and three difficulty levels for evaluating whether image generators can reproduce dense, prompt-specified text across four to twelve annotated regions, and reports results for 24 model configurations using a VLM judge.

What This Paper Is About

Text-to-image models have gotten good at rendering short strings, so the authors argue evaluation must now test whether that ability holds up under sustained load — hundreds to thousands of characters spread across many separate regions. UltraText Bench supplies every target string directly in the prompt, plus a separate structured reference describing each region's content, placement, and carrier, and then asks a VLM judge to score the whole generated image against that full reference. The goal is to measure text fidelity, text clarity, spatial quality, and scene quality separately, rather than reporting a single average that can hide failures.

Key Contributions

  1. UltraText Bench dataset: 432 prompts spanning 24 real-world scene categories in six domains, at three difficulty levels (L1 Hard, L2 Very Hard, L3 Extreme), split equally between English and Chinese, with exact target strings and a uniform per-region reference for prompt-only generation. Each prompt requests four to twelve text regions.
  2. Whole-scene evaluation protocol: A shared rubric applied with the Q-Judger model from Qwen-Image-Bench to the generated image and the complete reference, returning six image-level scores that map to four reporting dimensions, with evaluation coverage reported separately.
  3. Experimental analysis: Comparisons of 24 model configurations across dimensions (fidelity, clarity, spatial, scene), across workload levels and languages, and across Base/Turbo variant pairs.
  4. Human evaluation: Ten participants took part in a human evaluation of the automatic scores.

Main Findings

  • The benchmark is large and genuinely dense: 432 prompts contain 407,918 GT characters and 2,926 regions in total, with a mean of 944.25 GT characters and 6.77 regions per prompt overall (ranges 278–5,230 characters and 4–12 regions).
  • Fidelity and clarity diverge: Very clear lettering is not the same as correct content. Z-Image-Turbo receives 74.41 Clarity but 40.75 Fidelity; Qwen-Image-2512 receives 79.89 Clarity and 59.30 Fidelity. Both also score higher on scene quality than on fidelity.
  • Turbo variants trade fidelity for clarity: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. The paper states the Z-Image and LLaDA Turbo variants receive higher clarity but lower fidelity scores than their Base counterparts.
  • Performance collapses as workload rises: Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3.
  • Strong separation between open-weight and API-only models: The best open-weight model, Boogu-Image-0.1-Base, reaches a composite of 77.89, while API-only GPT Image 2 [Low] reaches 99.35 and Nano Banana 2 reaches 97.00; Seedream 5.0 Pro reaches 91.45, Qwen-Image-2.0-Pro 90.21, Qwen-Image-3.0 88.33, and WAN-2.7-Image-Pro 85.73. The lowest reported composite is SD3.5 Large at 7.48.
  • Averaging can hide failures: The paper's menu example receives TA = 10 and TC = 100, giving a Fidelity of 55; reporting only that average conceals the large gap between correctness and completeness. Grid positions are coarse (3×3) and several regions can share a cell, while layout quality rates organization without explicit GT for every within-cell relation.
  • The judge sometimes disagrees with the image: Qualitative examples retain disagreements with Q-Judger — target-related chat text is visible in one figure despite zero accuracy and completeness ratings, and another figure contains degraded small characters despite maximum ratings.
  • Quantitative human-agreement statistics are not reported: Ten participants took part in the human evaluation, but inter-rater and human–judge agreement statistics are stated as not reported. The paper also reports descriptive means without confidence intervals, and notes that close comparisons require uncertainty estimates across prompts.
  • Coverage can drop independently of quality: Failed evaluations remain unscored and reduce coverage rather than being counted as low-quality scores; image coverage is the fraction of planned images with valid scores and prompt coverage is the fraction of prompts with at least one valid score.

Methodology in Plain English

The authors built a prompt-only generation test. Each benchmark record pairs a natural-language scene description containing every target string with a separate structured annotation that the generator never sees. That annotation lists, for each of four to twelve regions, the exact target text, one of nine grid positions, a relative size, a text type, a carrier (the physical surface or digital element bearing the text), and an importance level. The reference specifies coarse spatial metadata only; the protocol returns image-level scores, and region-level attribution is named as a future direction.

The 24 categories are grouped into six domains — Signage & Labels (sign, label, poster, billboard), Documents & Print (article, newspaper, letter, resume), Commercial (menu, receipt, invoice, product packaging), Digital Interfaces (webpage, slide, social media, dashboard), Structured Data (schedule, form, certificate, code), and Creative & Special (caption, dialogue, comic panel, infographic). Each category × level × language cell contains three prompts. With four samples per prompt, a complete run yields 12 images per cell, 288 per level and language, 864 images per language, and 1,728 images per model configuration.

Difficulty is defined by language-specific target-character bands that include spaces and punctuation. English bands are 250–600, 600–1200, and 1200+ GT characters; Chinese bands are 350–600, 550–900, and 600+. Realized English means rise from 489.90 at L1 to 2,386.68 at L3, with mean region counts rising from 4.67 to 9.38; Chinese means rise from 394.62 to 743.62, with region counts from 4.93 to 8.40. The paper notes that because text load and region count change together, the levels do not isolate their individual effects, and it says L3 English is open ended and some Chinese bands overlap.

For evaluation, Q-Judger from Qwen-Image-Bench receives the image and the complete reference and returns six raw scores on a 0–100 scale: text accuracy (TA), text completeness (TC), text readability (TR), position correctness (PC), layout quality (LQ), and scene integration (SI). These map to four reporting dimensions — text fidelity = (TA + TC)/2, text clarity = TR, spatial quality = (PC + LQ)/2, and scene quality = SI — and a composite computed as 0.60 × fidelity + 0.30 × clarity + 0.05 × spatial + 0.05 × scene. The paper emphasizes these are VLM ratings, not measured percentages of correct characters or recovered regions. Scores are aggregated prompt-macro (each represented prompt weighted equally), and the bilingual result averages the EN and ZH prompt-macro means equally. All 432 prompts were manually reviewed and verified by automated checks for region identifiers, the six attributes and their allowed values, and character statistics.

Why This Matters

Impact on research. The paper argues that scoring text rendering by a single OCR-based number cannot distinguish spelling errors from missing text from unreadable text from misplaced text, and that benchmarks with different inputs — glyph controls, layout controls, or source images — cannot be placed on one scale. UltraText Bench keeps the model's input to a prompt alone and holds the reference fixed across scenes, languages, and levels, which lets researchers ask whether text rendering transfers across carriers and layouts as text load grows.

Real-world applications (all scene types named in the paper):

  • Signage and public information: signs, labels, posters, billboards, and opening-hours plates, where perspective and weathered surfaces complicate rendering.
  • Documents and print: articles, newspapers, letters, and resumes, where paragraph flow, columns, and headline hierarchy must hold.
  • Commercial paperwork and packaging: menus, receipts, invoices, and product packaging, where an incorrect total or a misspelled item changes the information conveyed.
  • Digital interfaces and creative layouts: webpages, slides, social media cards, dashboards, forms, certificates, code, captions, dialogue, comics, and infographics.

Industry relevance. The leaderboard separates open-weight and API-only models, which directly informs deployment choices for teams generating text-bearing images, and the workload breakdowns show that a model performing well at L1 may degrade substantially at L3 — Qwen-Image-2512's English composite drops from 86.50 to 42.86. The availability of a public repository (https://github.com/LINs-lab/UltraText_Bench) makes the suite reusable for model selection and regression testing.

Future Directions

  • Region-level attribution. The paper explicitly states that the reference specifies coarse spatial metadata and the protocol returns image-level scores, so region-level attribution remains a future direction. The output contains no region-level scores or transcriptions.
  • Judge reliability and uncertainty. The authors report descriptive means without confidence intervals and note that images from a shared prompt are related observations, so close comparisons need uncertainty estimates across prompts. Quantitative inter-rater and human–judge agreement statistics are not reported here.
  • Separating text load from region count. Because the three difficulty levels increase both GT-character load and region count together, the levels do not isolate the individual effect of either factor; the paper notes both change together while also providing region count as an additional difficulty axis.
  • Disagreements with the automatic judge. The qualitative examples retain cases where Q-Judger's ratings conflict with visible image content, leaving open how to detect and correct evaluator-side errors such as a VLM reader silently correcting characters the image renders wrongly, or OCR recognition errors.

Target Audience

Researchers and engineers working on text-to-image generation and visual text rendering, especially those building or training models that must produce documents, signage, packaging, and interfaces. It is also relevant to practitioners who need to choose between open-weight and API-only image models for text-heavy production work, and to evaluation researchers interested in rubric-based VLM judging, bilingual evaluation design, and the relationship between automatic scores and human judgment.

Authors’ abstract

Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.

Read the original paper