Research
Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
Overview Research area: Computer vision, specifically real-world scene text image super-resolution (SR) — restoring high-resolution images from degraded low-resolution inputs while keeping text readab
- arXiv
- 2510.21590
- Published
- 2025-10-24
- Authors
- Minxing Luo, Linlong Fan, Wang Qiushi, Ge Wu, Yiyan Luo, Yuhang Yu, Jinwei Chen, Yaxing Wang, Qingnan Fan, Jian Yang
AI summary
Overview
Research area: Computer vision, specifically real-world scene text image super-resolution (SR) — restoring high-resolution images from degraded low-resolution inputs while keeping text readable.
Technical level: Advanced. The paper builds on latent diffusion models, ControlNet-style conditioning, VAE latents, UNet denoisers, segmentation losses, and OCR-based evaluation.
Scope: The paper proposes a two-stage "text-first, image-later" framework (TiGeSR) that restores glyph structure in text regions before performing full-image super-resolution, and releases a new Chinese scene text benchmark (UZ-ST) captured at focal lengths from 14 mm to 200 mm with a maximum zoom of ×14.29.
What This Paper Is About
General image super-resolution methods can synthesize plausible natural textures but often destroy text, turning characters into gibberish — a problem that is worse for Chinese, where a single distorted or missing stroke changes meaning. Text-only super-resolution methods improve readability but, lacking global background constraints, introduce style inconsistencies and block artifacts between text and background. TiGeSR is designed to break this trade-off by explicitly decoupling glyph restoration from image enhancement: it reconstructs text structures first and uses them as conditional guidance for the full-image super-resolution stage.
Key Contributions
-
TiGeSR framework. The first two-stage scene text super-resolution framework with a "text-first, image-later" paradigm that decouples glyph restoration from image enhancement, aimed at improving both readability and visual quality.
-
UZ-ST dataset. The first scene text benchmark with a maximum zoom of ×14.29, containing aligned, richly annotated low-resolution/high-resolution pairs for challenging real-world conditions. It is described as the first Chinese scene text dataset with a highest zoom of ×14.29, and it covers Chinese characters, multi-line text, and real-world degradation greater than ×4 — the three properties the paper says TextZoom, CTR, and Real-CE each lack in some combination.
-
A two-phase training strategy for text restoration. Phase 1 trains on both synthetic and real data so the model captures real-world degradation patterns; Phase 2 freezes the RGB output block and the mask output block and trains the UNet only on synthetic data to refine mask quality, combining synthetic precision with real-world degradation.
-
A cascade coarse-to-fine alignment pipeline. Images are sorted by focal length and each low-quality image is sequentially aligned to its next higher-focal neighbor, then refined against the 200 mm ground truth, with manual filtering for problematic cases.
Main Findings
-
State-of-the-art on both benchmarks. On Real-CE, TiGeSR reaches PSNR 24.12, SSIM 0.839, LPIPS 0.164, DISTS 0.125, FID 38.72, and OCR-A 67.3%. On UZ-ST (average across difficulty levels) it reaches PSNR 25.48, SSIM 0.830, LPIPS 0.196, DISTS 0.156, FID 20.01, and OCR-A 43.0%. The authors state it achieves OCR-A scores above 0.67 and 0.43 respectively.
-
Closest competitor on text accuracy. TADiSR records OCR-A 64.7% on Real-CE (PSNR 23.83) and 36.6% on UZ-ST (PSNR 24.61), the second-highest OCR-A in both cases. The paper attributes TADiSR's limits to the resolution of its cross-attention mechanism, which struggles under severe degradation and on small text.
-
Only method with positive ΔOCR-A relative to the low-resolution input. On cropped text regions, TiGeSR achieves ΔOCR-A of +2.5% on Real-CE and +1.3% on UZ-ST. Every listed competitor is negative — for example Real-ESRGAN −8.8%/−4.7%, HAT −8.2%/−3.9%, SeeSR −27.4%/−18.9%, TADiSR −1.0%/−5.1%.
-
UZ-ST fine-tuning improves other architectures too. OSEDiff improves from OCR-A 22.8% to 28.9% (PSNR 23.50 to 25.07) and DiT4SR from 19.3% to 23.7% (PSNR 22.58 to 23.17) when trained with UZ-ST. The authors note these models lack a dedicated text-structure mechanism and therefore still cannot beat TiGeSR, whose own score rises from 40.0% to 43.0%.
-
Low dependence on the OCR module. On the 35 mm subset of UZ-ST, replacing the OCR-predicted text with null text yields OCR-A 40.4% and random text yields 40.3%, both above TADiSR's 35.5%, while predicted text gives 44.6%. A separate ablation shows performance degrading gracefully as OCR randomness increases: 100% random 40.3%, 50% random 40.6%, 20% random 43.7%, 0% random 44.6%.
-
Ablation on the text mask source (Real-CE, stage 2 fixed). TiGeSR's mask gives OCR-A 67.3%, versus empty mask 59.5%, standard font guidance 55.3%, SAM-TS extraction 57.9%, and LDM guidance 60.1%. TiGeSR also has the best LPIPS (0.164), DISTS (0.125), and FID (38.72) in this table.
-
Per-zoom behavior on UZ-ST. TiGeSR scores OCR-A 63.2% at 85 mm (×2.35), 44.6% at 35 mm (×5.71), and 16.0% at 14 mm (×14.29), showing accuracy falls sharply at the most extreme zoom.
-
Alignment pipeline matters. The cascade coarse-to-fine alignment achieves PSNR 23.00, MSE 441.52, SSIM 0.7279, NCC 0.9184, and AKD 241.44, versus RealSR (PSNR 13.87, NCC 0.3285) and single-pass SIFT alignment (PSNR 12.77, NCC 0.4313).
-
Efficiency. TiGeSR Stage 1 uses 6215.85 GFLOPs at 663.5 ms; Stage 2 uses 3734.01 GFLOPs at 360.06 ms. For comparison, TADiSR uses 4497.96 GFLOPs at 342.32 ms, HAT 6670.32 GFLOPs at 1086.68 ms, DiffTSR 58502.14 GFLOPs at 8610.59 ms, DiT4SR (w/o llava) 160787.14 GFLOPs at 17385.14 ms, and DreamClear (w/o llava) 412843.67 GFLOPs at 83193.81 ms.
-
Qualitative failure modes of prior work. DiffTSR+HAT is reported to produce text colors inconsistent with the original LR, and in one UZ-ST example to distort a face near the word "Ory". MARCONet and DiffTSR struggle with long text sequences and large width-to-height aspect ratios; MARCONet additionally fails on complex scripts due to its primitive architecture.
Methodology in Plain English
The framework has two sequential stages.
Stage 1 — Text Restoration. An OCR detector first finds text regions in the low-resolution input and extracts their text content as a semantic condition. Each cropped region is encoded by a VAE into a latent, concatenated with noise, and iteratively denoised by a UNet that produces two branches: an RGB branch for appearance and a mask branch for structure. The text content is embedded and fused into the denoising via cross-attention to guide structure recovery. After T denoising steps, the mask branch output is decoded and all restored regions are reassembled at their original positions to form a text mask. Training uses a combined objective of a text-control diffusion loss and a segmentation-oriented loss built from MSE, Focal, and Dice terms, with the two-phase synthetic/real training schedule described above.
Stage 2 — Image Enhancement. A ControlNet-like network takes the low-resolution latent and the reconstructed text mask latent and denoises at a specified timestep using a null-text embedding, producing the super-resolved latent. Training combines MSE and LPIPS reconstruction losses with an edge loss computed by applying Sobel operators to the ground-truth and predicted high-resolution images.
Data and evaluation. UZ-ST was captured with a ViVO X200 Ultra at four fixed focal lengths (14 mm, 35 mm, 85 mm, 200 mm), with images at 200 mm treated as ground truth. It contains 5,036 image pairs and 49,675 text lines, covering street views, book covers, advertisements, menus, posters, and documents under daylight, indoor, and nighttime lighting. The pairs split into 1,439 at ×14.29, 1,798 at ×5.71, and 1,799 at ×2.35 zoom modes, with 470, 589, and 581 pairs randomly selected for evaluation in each mode respectively. Per subset, the 14 mm set has 1,439 images, 15,073 lines, 10.47 mean lines per image, and 754 images with more than 5 lines; the 35 mm set has 1,798 images, 17,263 lines, 9.60, and 867; the 85 mm set has 1,799 images, 17,339 lines, 9.64, and 869. Full-image quality is measured with PSNR, SSIM, LPIPS, DISTS, and FID; cropped text regions are measured with PSNR_cr, SSIM_cr, LPIPS_cr, and DISTS_cr; and text accuracy uses OCR-A, a Levenshtein-ratio-based comparison between OCR output and ground-truth annotations. Training mixes synthetic data built on LSDIR with text rendered via LBTS and Real-ESRGAN degradation, plus real data from Real-CE and UZ-ST; misaligned Real-CE pairs were filtered and reannotated to yield 337 training and 188 testing pairs. Stage 1 uses an IDM-based architecture, Stage 2 is based on Stable Diffusion 3.5 with a tile-based inference strategy, and the second stage is pretrained on synthetic data and fine-tuned on real data.
Why This Matters
Impact on research. The paper reframes scene text super-resolution as a decoupling problem rather than a single generative-prior problem, arguing that text and non-text regions need different treatment. It also argues that existing benchmarks capped at mild degradation (≤×4) fail to reflect real-world long-distance capture, and supplies a dataset at up to ×14.29 zoom to test that gap.
Real-world applications:
- Restoring shop signs, posters, menus, and documents captured from a distance or at low zoom.
- Enhancing accessibility for low-vision users who need to read text in photographs.
- Document restoration and archiving where degraded characters must remain semantically correct.
- Navigation assistance, where signage text must be read correctly in real time.
Industry relevance. The authors are affiliated with vivo BlueImage Lab and vivo Mobile Communication, and the dataset was captured with a ViVO X200 Ultra using its four fixed focal lengths. The paper's efficiency table is directly relevant to mobile deployment: TiGeSR's Stage 2 runs at 3734.01 GFLOPs and 360.06 ms, well below several diffusion baselines. The ethics statement also flags misuse risks such as recovering text from personal or social media images or surveillance deployment, recommending privacy-preserving measures such as watermarking or selective inpainting. The paper states that code and dataset will be released.
Future Directions
- Closing the extreme-zoom gap. OCR-A falls to 16.0% at ×14.29 on UZ-ST, far below the 63.2% at ×2.35, so the hardest regime of the new benchmark remains largely unsolved.
- Reducing reliance on OCR at inference. The ablations show graceful degradation with null, random, and partially random OCR text, but also show a measurable drop from 44.6% to 40.3% OCR-A on the 35 mm subset, leaving room for structure supervision that needs no text prediction at all.
- Extending beyond Chinese scene text. The paper's motivation centers on the stroke sensitivity of Chinese characters; whether the same two-stage decoupling transfers to other complex scripts is not reported.
- Efficiency and on-device deployment. Two stages still incur combined cost (6215.85 GFLOPs at 663.5 ms plus 3734.01 GFLOPs at 360.06 ms), and the paper does not report end-to-end latency for the full pipeline on mobile hardware.
Target Audience
Researchers and engineers working on image super-resolution, diffusion-based generative restoration, OCR and document analysis, and scene text understanding — particularly those dealing with Chinese text or mobile camera pipelines. It is also useful for dataset builders interested in alignment strategies for large-zoom-factor paired data, and for practitioners who need to decide between full-image SR and text-region SR for a production system.
Authors’ abstract
Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided supEr-Resolution), a novel two-stage framework that breaks this trade-off through a "text-first, image-later" paradigm. TIGER explicitly decouples glyph restoration from image enhancement: it first reconstructs precise text structures and uses them to guide full-image super-resolution. This ensures high fidelity and readability. To support comprehensive training and evaluation, we present the UZ-ST (UltraZoom-Scene Text) dataset, the first Chinese scene text dataset with extreme zoom. Extensive experiments show TIGER achieves state-of-the-art performance, enhancing readability and image quality.