Skip to content
AI.info

Research

UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images

UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images Overview Research area: Computer vision and generative AI for graphic design — specifically diffusion/Di

arXiv
2512.20479
Published
2025-12-23
Authors
Yiming Zhao, Yuanpeng Gao, Yuxuan Luo, Jiwei Duan, Shisong Lin, Longfei Xiong, Zhouhui Lian

AI summary

UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images

Overview

Research area: Computer vision and generative AI for graphic design — specifically diffusion/DiT-based rendering, editing, and synthesis of stylized text (typography) inside design images such as posters, banners, and advertisements.

Technical level: Advanced. The paper assumes familiarity with diffusion transformers (DiTs), Rectified Flow, VAE latents, LoRA fine-tuning, DPO/RLHF for diffusion models, GRPO, and multimodal LLM encoders.

Scope in one sentence: The paper proposes UTDesign, a single DiT-based framework that both edits existing stylized text and generates new stylized text in design images for English and Chinese, and wires it together with a text-to-image model and an MLLM layout planner into a full text-to-design pipeline.

(Citation: arXiv:2512.20479v1 [cs.CV], 23 Dec 2025; published as a SIGGRAPH Asia 2025 Conference Paper, December 15–18, 2025, Hong Kong; DOI 10.1145/3757377.3763923; code at https://github.com/ZYM-PKU/UTDesign. Authors are affiliated with the Wangxuan Institute of Computer Technology at Peking University, Kingsoft Office, and the State Key Laboratory of General Artificial Intelligence.)

What This Paper Is About

Diffusion text-to-image models generate attractive design images but render text poorly, especially small typography and non-Latin scripts like Chinese, and they rarely let a user make fine-grained edits while keeping the original font style and texture intact. UTDesign tackles both problems with one model: it edits stylized text in an existing design so the replacement keeps the reference glyph style, and it generates new stylized text conditioned on a background image, a prompt, and a layout. The framework is then extended into an end-to-end "text-to-design" pipeline that turns a user description into a finished design image with accurately rendered Chinese and English text.

Key Contributions

  1. A unified DiT-based framework for stylized text editing and generation in design images. Equipped with a customized transparency glyph VAE decoder, the system outputs high-quality transparent RGBA text foregrounds, and the authors report state-of-the-art performance in both text rendering accuracy and stylistic consistency.
  2. Two new datasets: a large-scale synthetic dataset (SynthGlyph) for training text editing models, and a curated real-world design image dataset (DesignText) with fine-grained text annotations that is extendable for training layout planning and text synthesis models.
  3. An integration into a fully automated text-to-design (T2D) system by combining UTDesign with pre-trained T2I models and an MLLM-based layout planner, translating user intentions into finalized graphic designs.
  4. A new evaluation benchmark (UTDesign-Bench) with 1,000 cases split into an editing track and a generation track, plus a two-stage layout planner whose MLLM is further optimized with GRPO using rule-based rewards.

Main Findings

  • Editing results (UTDesign-Bench-Edit): UTDesign scores best on every reported column against DiffUTE, AnyText-Edit, and AnyText2-Edit — FID 10.81, LPIPS 0.0883, CLIP-Sim 0.8222, Precision 0.9568, Recall 0.9482, F-Score 0.9518, NED 0.0612, Accuracy 0.8370. For comparison, the next-best FID is AnyText2-Edit at 20.68 and the next-best Accuracy is AnyText-Edit at 0.3538.
  • Generation results (UTDesign-Bench-Gen): Against Kolors, Cogview4, BrushYourText, AnyText-Gen, AnyText2-Gen, Glyph-ByT5-v2, PosterMaker, Seedream 3.0, and GPT-4o, UTDesign reports FID 72.07 (best), F-Score 0.8716 (best), NED 0.1590 (best), and Accuracy 0.6840 (best, versus GPT-4o's 0.5772 and Seedream 3.0's 0.4885). It is not best on LPIPS (0.6973, versus Seedream 3.0 at 0.6903) or CLIP-Sim (0.2609, versus GPT-4o at 0.2710).
  • Precision/recall trade-off in generation: Glyph-ByT5-v2 has the highest Precision (0.9066 vs UTDesign's 0.8784) and Seedream 3.0 the highest Recall (0.9603 vs UTDesign's 0.9106), but UTDesign leads on the combined F-Score.
  • Open-source T2I models are not competitive on this task: The paper notes that although open-source T2I models show relatively high CLIP similarity, their overall performance is suboptimal due to low aesthetic quality and frequent omission of visual text (Kolors Accuracy 0.0010, CogView4 Accuracy 0.0600).
  • Comparison to proprietary systems: The authors state UTDesign achieves performance on par with closed-source systems and significantly outperforms all open-source baselines, with native support for transparent RGBA text foregrounds.
  • Post-training ablation (Table 2): Progressing from the aligned base model to +SFT to +DPO improves NED (0.2770 → 0.2624 → 0.2604) and Accuracy (0.5570 → 0.5740 → 0.5930), while PickScore rises from 19.83 to 19.78 to 20.10. FID moves only slightly (73.58 → 73.12 → 73.11). DPO further enhances both aesthetic appeal and diversity.
  • Layout planner ablation (Table 3): Pretrained MLLM is insufficient (coarse R_iou 0.0462, FID 23.77). Adding SFT improves coarse R_iou to 0.3637 and lowers FID to 10.05; adding GRPO raises coarse R_iou further to 0.5219 and lowers FID to 8.146. Fine-grained R_iou rises 0.0866 → 0.5883 → 0.6873, and the balance penalty −R_bl falls from 0.5435 → 0.1781 → 0.0912.
  • User study: Because FID biases and OCR inaccuracy mean the quantitative metrics do not fully reflect human preferences, the authors ran a user study against Seedream 3.0 and GPT-4o on prompt matching, overall aesthetics, and text accuracy. UTDesign showed comparable overall image quality and a clear advantage in text rendering accuracy.

Methodology in Plain English

The core design decision is that instead of asking one model to learn everything at once, the authors train in three progressive stages so the model first separates what a glyph says from how it looks.

  • Data first. They built a synthetic glyph dataset from 4,194 TrueType fonts and 6,857 different characters per font (Chinese glyphs from the GB6763 standard plus 94 English letters and symbols), producing roughly 28.8M stylized character instances, with extra color and texture augmentation. From this they sample triplets of a content reference, a style reference, and an RGBA ground truth. They also curated 115.5k real design samples, each with the design image, an extracted background, a caption, text content with glyph-level boxes, and an RGBA text foreground.
  • Stage 1 — learn editing. A DiT backbone is trained from scratch with two separate encoders: a DINOv2-based ViT for glyph content and a CLIP-initialized encoder for glyph style, each projected into a shared latent space. Training follows the Rectified Flow objective, predicting a velocity field from noisy latents conditioned on both content and style references. The content and style embeddings are concatenated with the noisy latents and processed by fusion DiT blocks (full attention plus a tanh gate controlling style strength) and single DiT blocks (parallel self-attention/feedforward for efficiency), with 3D-RoPE to distinguish token type and character identity, and RMSNorm throughout. Style references are perturbed with Gaussian blur, down-sampling, noise, and background changes to improve generalization. The model is then fine-tuned on the real design dataset.
  • Stage 2 — connect the encoders. A frozen MLLM reads the design background plus the textual description and emits hidden states; a Perceiver Resampler converts these variable-length outputs into a fixed-size representation. Only the resampler is trained, using an L2 loss to match the style embedding space learned in Stage 1.
  • Stage 3 — turn editing into generation. The style encoder is swapped out for the trained MLLM encoder. The model is fine-tuned with LoRA under a Rectified Flow loss on a filtered high-quality subset, then optimized with Diffusion-DPO: multiple candidate outputs are scored and ranked by a pre-trained aesthetic reward model to form win–lose pairs, and the reference model in the DPO objective is the Stage-1 editing model.
  • Transparent output. A transparency glyph VAE uses a pretrained FLUX VAE encoder with a decoder extended by extra convolution layers to produce an alpha channel, trained with L2 plus LPIPS losses on alpha-blended glyphs. This yields RGBA foregrounds that can be composited onto any background.
  • Layout planning. A two-stage MLLM planner predicts line-level boxes first and then per-character positions within those lines — important for Chinese, which is sensitive to character-level layout. It is trained with SFT on real layouts and then with GRPO using three rule-based rewards: mean IoU against ground truth, a penalty for overlapping boxes, and a penalty for uneven box sizes.
  • Inference. For editing, the style is taken from a user-specified region, the DiT produces a new foreground, an inpainting model clears the original text, and the layout planner places the new text. For generation, the DiT produces foreground glyphs from the background and description, the planner handles line-level and character-level typography, and everything is composited. Adding a pre-trained T2I model (e.g., FLUX) to synthesize the background gives the full text-to-design pipeline.

Why This Matters

  • Research impact: The paper argues that prior UNet-based text-rendering methods improve text accuracy but often compromise overall aesthetic quality, and that there is a lack of open-source models matching proprietary state-of-the-art text rendering. UTDesign contributes an open framework, two datasets (SynthGlyph, DesignText), a benchmark (UTDesign-Bench), and a transparency-aware VAE design that produces editable RGBA text layers rather than flattened pixels.
  • Real-world applications:
    • Automated poster, banner, and advertisement creation from a text description.
    • Editing existing marketing material — swapping headline copy while preserving the original font style and texture.
    • Multilingual design workflows requiring both English and Chinese typography.
    • Layer-based design tools where generated RGBA text foregrounds can be moved, restyled, or repositioned by a human designer afterward.
  • Industry relevance: The work is a collaboration with Kingsoft Office (WPS), pointing directly at office/design software; the second author team and funding context suggest the pipeline targets production design tools. Its reported ability to approach commercial tools (GPT-4o, Seedream 3.0) while remaining open-source is the central practical claim.

Future Directions

  • Closing the remaining gap to proprietary systems. UTDesign matches but does not clearly beat commercial systems such as GPT-4o and Seedream 3.0, which still lead on some individual metrics (LPIPS, CLIP-Sim, Precision, Recall).
  • Better evaluation. The authors explicitly state that FID bias and OCR limitations mean the metrics do not adequately reflect human preferences, and that OCR-based scores diverge from actual performance — better, human-aligned benchmarks are an open problem.
  • Extending language coverage. The paper supports English and Chinese; it notes other systems (e.g., ART) fail on non-Latin languages due to base-model limits and that non-Latin scripts remain a general weakness in text rendering.
  • Further layout and RL optimization. The GRPO-based layout planner ablations show the largest jumps from SFT, with GRPO adding further gains, suggesting more reward design or planner training is a productive direction. The paper's ablation discussion is truncated in the provided content after noting that GRPO further enhances performance, so the complete conclusion of that analysis is not reported here.

Target Audience

Researchers and engineers working on diffusion-based image generation, text rendering in images, and AI-assisted graphic design, as well as applied teams building design-automation products that need precise multilingual typography. It is most useful to readers already comfortable with DiT architectures, Rectified Flow, and preference-optimization methods for diffusion models; the transcribed paper content is heavily equation- and table-driven and assumes that background.

Authors’ abstract

AI-assisted graphic design has emerged as a powerful tool for automating the creation and editing of design elements such as posters, banners, and advertisements. While diffusion-based text-to-image models have demonstrated strong capabilities in visual content generation, their text rendering performance, particularly for small-scale typography and non-Latin scripts, remains limited. In this paper, we propose UTDesign, a unified framework for high-precision stylized text editing and conditional text generation in design images, supporting both English and Chinese scripts. Our framework introduces a novel DiT-based text style transfer model trained from scratch on a synthetic dataset, capable of generating transparent RGBA text foregrounds that preserve the style of reference glyphs. We further extend this model into a conditional text generation framework by training a multi-modal condition encoder on a curated dataset with detailed text annotations, enabling accurate, style-consistent text synthesis conditioned on background images, prompts, and layout specifications. Finally, we integrate our approach into a fully automated text-to-design (T2D) pipeline by incorporating pre-trained text-to-image (T2I) models and an MLLM-based layout planner. Extensive experiments demonstrate that UTDesign achieves state-of-the-art performance among open-source methods in terms of stylistic consistency and text accuracy, and also exhibits unique advantages compared to proprietary commercial approaches. Code and data for this paper are available at https://github.com/ZYM-PKU/UTDesign.

Read the original paper