Research
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
Overview Research area: Computer vision — image super-resolution (SR), specifically scene-text super-resolution using vision-language-guided latent diffusion models. Technical level: Advanced. The pap

- arXiv
- 2510.26339
- Published
- 2025-10-30
- Authors
- Mingyu Sung, Seungjae Ham, Kangwoo Kim, Yeokyoung Yoon, Sangseok Yun, Il-Min Kim, Jae-Mo Kang
AI summary
Overview
Research area: Computer vision — image super-resolution (SR), specifically scene-text super-resolution using vision-language-guided latent diffusion models.
Technical level: Advanced. The paper assumes familiarity with latent diffusion models, ControlNet-style conditioning, classifier-free guidance, and OCR evaluation metrics.
Scope: The paper proposes GLYPH-SR, a VLM-guided latent diffusion framework that jointly optimizes perceptual image quality and character-level text legibility in scene-text super-resolution, evaluated on SVT, SCUT-CTW1500, and CUTE80 at ×4 and ×8 scales.
What This Paper Is About
Standard super-resolution research is tuned to global metrics like PSNR/SSIM or learned perceptual scores (LPIPS, MANIQA, CLIP-IQA, MUSIQ), which are dominated by large image areas and therefore barely penalize corruption in small text regions — the paper notes text regions are often well below 1% of the image. As a result, SR models either hallucinate sharp but incorrect characters or conservatively preserve blurry input, and both failure modes break downstream optical character recognition (OCR) even when the rest of the image looks sharp. GLYPH-SR reframes scene-text SR as a bi-objective problem — optimizing visual realism and text legibility together — rather than treating text as generic high-frequency texture.
Key Contributions
- Bi-objective formulation and dual-axis evaluation. The authors explicitly cast SR in text-rich scenes as joint optimization of image quality and readability, and standardize a protocol that reports perceptual SR metrics (MANIQA, CLIP-IQA, MUSIQ) alongside OCR-aware measures (word/character accuracy, edit distance, F1) so that small text regions are not underweighted.
- Text-SR Fusion ControlNet (TS-ControlNet) with time-balanced guidance. A dual-branch ControlNet fuses token-level OCR strings with verbalized locations (𝒮_TXT) and a scene caption (𝒮_IMG). The SR branch is frozen while the text branch is fine-tuned; residual mixing injects complementary cues into the latent diffusion model without disrupting its generative prior. A lightweight "ping-pong" scheduler λ_t alternates text-centric and image-centric conditioning along the denoising trajectory, modulating both embedding fusion and residual injection.
- Factorized synthetic corpus and comprehensive validation. A four-partition synthetic corpus independently perturbs glyph quality and global image quality, enabling targeted text restoration while the SR branch stays frozen.
- Released artifacts. Code, pretrained models, data-generation scripts, and an evaluation suite are released to support reproducibility.
Main Findings
- OCR F1 gains up to +15.18 percentage points. Across SVT, SCUT-CTW1500, and CUTE80 at ×4 and ×8, GLYPH-SR improves OCR F1 by up to +15.18 pp over diffusion/GAN baselines (SVT ×8, OpenOCR), while maintaining competitive MANIQA, CLIP-IQA, and MUSIQ.
- Best OpenOCR F1 in 5 of 6 settings, best GOT-OCR F1 in 4 of 6. On SVT ×8 the method is best across all six reported metrics; on CUTE80 ×8 it leads all SR metrics while also securing the top OpenOCR F1 score of 63.66. On the most challenging benchmarks the paper reports surpassing competitors by a large margin (e.g., +12.0 pp on CUTE80 ×8).
- Perceptual quality stays competitive. GLYPH-SR ranks first or second in 26 out of 30 test cases across MANIQA, CLIP-IQA, and MUSIQ, frequently outperforming other diffusion models such as DiffBIR and SUPIR.
- The baseline trade-off is documented concretely. On SVT ×4, DiffBIR achieves strong SR metrics (47.82 MANIQA / 71.18 MUSIQ) but suffers text hallucination, yielding a low OpenOCR F1 of 38.73. Conversely, StableSR attains a high LLaVA-NeXT F1 (73.91) through conservative restoration but a poor MUSIQ of 24.44.
- Robustness widens at ×8. The performance gap grows at ×8 scale, where GLYPH-SR avoids both the textual hallucination of GANs and the over-smoothing of generic diffusion methods.
- Both guidance types are necessary (ablation). Using OCR string + position gives clean reconstructions. Text-only guidance produces irregular kerning and warped baselines (e.g., "STASHOES COFFEE"). Position-only guidance yields partial or incorrect spellings ("STABHOUES SOFFCE"). Removing both produces the worst outcomes — severe hallucinations and geometric distortions reminiscent of generic diffusion SR.
- Binary ping-pong beats continuous ramps. Continuous schedules λ_t = g(σ_t) (e.g., noise-level monotone schedules) were also tested, but the square-wave "ping-pong" yielded the best OCR F1 at similar perceptual quality (default toggle period τ = 1).
Methodology in Plain English
The researchers start from a pretrained latent diffusion model — they adopt Juggernaut-XL as the backbone — and add a two-branch control module. One branch (SR-ControlNet) is a frozen, off-the-shelf super-resolution controller that handles overall image structure from the low-resolution input. The other branch (Text-ControlNet) is trainable and handles text rendering, driven by OCR-derived text-and-position pairs. An OCR module detects K text instances and returns position-text pairs, which are converted into structured natural-language prompts (for example, "HSBC is displayed at the center of the image"). A separate scene caption summarizes global attributes such as illumination, composition, and depth-of-field. During training the LDM backbone and the SR branch stay frozen and only the text branch is updated, so the model gains text-fixing ability without losing its generative prior. The two branches' residual hierarchies are blended before injection. A scheduler λ_t alternates between text-centric (λ_t = 0) and image-centric (λ_t = 1) guidance in a square wave, so text-focused phases inject precise glyph cues while image-focused phases stabilize global structure. To train this, the authors build a synthetic corpus of four mutually exclusive subsets — positive/negative text crossed with high/low image quality, denoted I^pos_HQ, I^pos_LQ, I^neg_HQ, I^neg_LQ — generated from the same raw text, with the text-position pairs always extracted from the positive-text, high-quality dataset so the model is explicitly told when incorrect text has been generated. Synthetic data were generated with LLaVA-NeXT, Nunchaku, and SUPIR. Sampling uses an Elucidated Diffusion Model (EDM) sampler with classifier-free guidance.
Why This Matters
Impact on research: The paper argues that existing SR metrics are structurally blind to text errors because they aggregate quality globally and are dominated by area, and it supplies a dual-axis evaluation protocol plus a factorized synthetic corpus that lets researchers disentangle text legibility from holistic perceptual quality. This reframing could shift how scene-text SR is benchmarked.
Real-world applications:
- Surveillance and autonomous driving, where signs and license plates must be read correctly for downstream decisions.
- Document analysis and OCR pipelines, where character hallucination silently corrupts extracted text.
- Retail analytics, including product labels and storefront signage reading.
- Any sign-reading or text-in-the-wild system that depends on recognizable characters rather than merely sharp-looking images.
Industry relevance: Systems that combine SR with OCR — document digitization, retail shelf monitoring, navigation, and accessibility tooling — benefit directly, because a sharp-looking but misread sign produces a downstream failure that is invisible to standard quality metrics.
Future Directions
- Explore multilingual scripts, since the current evaluation is on scene-text benchmarks with particular script coverage.
- Incorporate stronger geometric priors to further anchor glyph layout and kerning.
- Pursue tighter integration with end-to-end recognition systems, rather than treating OCR as an upstream prompt source.
- The paper's own ablation shows that location cues without semantics produce character-level ambiguity and semantics without location produce warped layouts, leaving open how best to represent text geometry beyond bounding-box position prompts.
- The training relies on a synthetic corpus with four quality partitions; whether this transfers to real degradation distributions is not reported in the provided content.
Target Audience
Researchers and engineers working on super-resolution, diffusion models, ControlNet-style conditioning, OCR, and document or scene-text understanding. It is also relevant to practitioners building production pipelines where image enhancement feeds a recognition stage, and to anyone designing evaluation protocols for perception systems that must balance visual realism against semantic correctness. The paper is advanced — it assumes comfort with latent diffusion, classifier-free guidance, and OCR metrics.
Authors’ abstract
Image super-resolution(SR) is fundamental to many vision system-from surveillance and autonomy to document analysis and retail analytics-because recovering high-frequency details, especially scene-text, enables reliable downstream perception. Scene-text, i.e., text embedded in natural images such as signs, product labels, and storefronts, often carries the most actionable information; when characters are blurred or hallucinated, optical character recognition(OCR) and subsequent decisions fail even if the rest of the image appears sharp. Yet previous SR research has often been tuned to distortion (PSNR/SSIM) or learned perceptual metrics (LIPIS, MANIQA, CLIP-IQA, MUSIQ) that are largely insensitive to character-level errors. Furthermore, studies that do address text SR often focus on simplified benchmarks with isolated characters, overlooking the challenges of text within complex natural scenes. As a result, scene-text is effectively treated as generic texture. For SR to be effective in practical deployments, it is therefore essential to explicitly optimize for both text legibility and perceptual quality. We present GLYPH-SR, a vision-language-guided diffusion framework that aims to achieve both objectives jointly. GLYPH-SR utilizes a Text-SR Fusion ControlNet(TS-ControlNet) guided by OCR data, and a ping-pong scheduler that alternates between text- and scene-centric guidance. To enable targeted text restoration, we train these components on a synthetic corpus while keeping the main SR branch frozen. Across SVT, SCUT-CTW1500, and CUTE80 at x4, and x8, GLYPH-SR improves OCR F1 by up to +15.18 percentage points over diffusion/GAN baseline (SVT x8, OpenOCR) while maintaining competitive MANIQA, CLIP-IQA, and MUSIQ. GLYPH-SR is designed to satisfy both objectives simultaneously-high readability and high visual realism-delivering SR that looks right and reds right.