Research
STELLAR: Scene Text Editor for Low-Resource Languages and Real-World Data
Overview Research area: Computer Vision — Scene Text Editing (STE), specifically diffusion-based image editing for multilingual scene text. Technical level: Intermediate. The core ideas are explained

- arXiv
- 2511.09977
- Published
- 2025-11-13
- Authors
- Yongdeuk Seo, Hyun-seok Min, Sungchul Choi
AI summary
Overview
- Research area: Computer Vision — Scene Text Editing (STE), specifically diffusion-based image editing for multilingual scene text.
- Technical level: Intermediate. The core ideas are explained plainly in the paper, but the method builds on diffusion models, OCR recognizers, and disentangled style/glyph encoders.
- Scope: The paper proposes STELLAR, a scene text editor that supports low-resource languages and real-world images, and introduces a paired dataset (STIPLAR) and a style-preservation metric (TAS).
What This Paper Is About
Scene text editing means changing the words in a photograph — a menu, a storefront sign, packaging — while keeping the original font, color, and background looking untouched. Most existing methods are trained on synthetic data and are dominated by English and Chinese, so they fail on languages with complex or less-represented scripts, and they degrade on real photos because of lighting, texture, and noise. STELLAR aims to edit text reliably in Korean, Arabic, and Japanese on real-world images, and to provide a better way to measure whether the visual style was actually preserved.
Key Contributions
- Language-adaptive glyph encoder. A lightweight transformer glyph encoder trained with language-specific recognizers instead of a single English-oriented one, so it can capture script-dependent structure such as right-to-left writing and context-dependent letter forms in Arabic.
- STIPLAR dataset. A real-world image-pair dataset for low-resource languages containing 9,764 Korean, 6,316 Arabic, and 2,023 Japanese pairs (18K total), built from MLT-2019 crops and Creative Commons web-crawled images, split 8:2 into training and evaluation sets.
- Text Appearance Similarity (TAS). A new metric that independently scores color, font, and background similarity and averages them, so style preservation can be evaluated even when no ground truth image exists.
- Multi-stage training strategy. Pre-training on synthetic text image pairs followed by fine-tuning on real-world pairs, which closes the synthetic-to-real domain gap without requiring a post-hoc inference trick.
Main Findings
- Best overall results in Korean and Arabic. On the STIPLAR evaluation set, STELLAR achieves the best overall performance in Korean and Arabic and remains competitive in Japanese. In Korean it reaches SSIM 0.5061, PSNR 16.1514, MSE 0.0301, and TAS 0.8596, versus TextFlux at SSIM 0.3409, PSNR 14.5132, MSE 0.0457, and TAS 0.8464.
- Recognition accuracy is the largest gain. STELLAR achieves the highest recognition accuracy across all languages. For Korean it records Rec.Acc 0.8042 and NED 0.9115, outperforming TextFlux by absolute margins of 0.5829 and 0.4279 respectively. Arabic reaches Rec.Acc 0.6840 and NED 0.8985, compared with 0.0714 and 0.4449 for TextFlux.
- Japanese is the weak spot. TextFlux slightly outperforms STELLAR in image quality and style preservation in Japanese (SSIM 0.4105 vs 0.3520, TAS 0.7857 vs 0.7714), which the authors attribute to the visual similarity between Japanese kanji and the Chinese characters prevalent in the baselines' training data. STELLAR still leads in recognition accuracy (0.4338 vs 0.4156).
- Average TAS improvement of 2.2%. Across languages, STELLAR achieves an average TAS improvement of 2.2% over the baselines.
- Stage 2 fine-tuning drives domain adaptation. Training on Stage 1 alone degrades image quality, TAS, and recognition accuracy in all languages (Korean TAS 0.8362 and Rec.Acc 0.6676 with Stage 1 only, versus 0.8596 and 0.8042 for the full model).
- Fine-tuning beats post-hoc adjustment. Applying the post-hoc technique (GaMuSa from TextCtrl) to a Stage 1 model still scores lower than STELLAR across all languages (Korean TAS 0.8375, Rec.Acc 0.6710).
- More real data helps. Downsampling the Stage 2 Korean and Arabic data to the Japanese dataset size still beats TextFlux reported in Table 2 but underperforms STELLAR in most metrics (Korean TAS 0.8543, Rec.Acc 0.7710).
- TAS is robust to text changes and sensitive to style changes. On synthetic Korean sets where only text content differed, SSIM and PSNR were low (0.5768 and 17.2873) while TAS was highest (0.8933); on color-modified sets, conventional metrics reported high similarity (SSIM 0.7783, PSNR 21.0446, MSE 0.0230, FID 21.2495) while TAS was comparatively lower at 0.8379.
- TAS works without ground truth. Comparing source versus generated images, TAS was comparable to or higher than ground-truth-based comparison in all languages (Korean 0.8608 w/ source vs 0.8641 w/ GT; Arabic 0.8726 vs 0.8636; Japanese 0.8031 vs 0.7627), despite lower SSIM, PSNR, MSE, and FID.
- Language-specific training helps from Stage 1. Even with Stage 1 training only, the model maintains higher recognition accuracy in Korean and Arabic than the baselines, which are primarily trained on Chinese characters.
- Synthetic-to-real efficiency. Stage 2 uses less than 5% of the synthetic data and 10% of the training epochs, yet rapidly adapts the model to real-world domains.
Methodology in Plain English
STELLAR builds on a diffusion-based baseline that separates visual style from text content. The pipeline has a glyph encoder, a style encoder, and a diffusion generator.
The glyph encoder turns the target text string into character-level features that guide text rendering through cross-attention in the generator. The distinctive change is that instead of one recognizer pre-trained mainly on English, STELLAR trains separate encoders using PPOCRv4 recognizers specialized for Korean, Arabic, and Japanese, so the glyph features reflect each script's structure.
The style encoder is a ViT-based module trained with four tasks at once — transferring text color, transferring font, removing text to reconstruct the background, and segmenting the text region mask. Its style features are injected into the generator through skip connections and middle blocks, encoding color, font, background, and spatial layout.
Training happens in two stages. Stage 1 pre-trains the generator on 200k synthetic text image pairs per language using only OCR-correct pairs, using the Stable Diffusion v1.5 checkpoint, images resized to 256x256, and a maximum target text length of 12. Stage 2 fine-tunes on the real-world STIPLAR pairs. Training used 2 NVIDIA H100 80GB GPUs, a learning rate of 1e-5, 100 epochs for Stage 1 (66 hours) and 10 epochs for Stage 2 (0.3 hours).
For evaluation, the authors compare against three mask-and-inpaint baselines that support multilingual editing — AnyText, AnyText2, and TextFlux — using the full image plus a mask (and a text prompt for AnyText and AnyText2), cropping a fixed-size patch up to 1024x1024 around the target text. Evaluation uses SSIM, PSNR, MSE, FID, TAS, OCR recognition accuracy, and Normalized Edit Distance. Inference uses 50 denoising steps, classifier-free guidance scale 2.0, and random seed 42.
TAS is computed by running both images through the style encoder. Color similarity renders a grayscale image of one text string and colorizes it with each image's texture features, then compares using normalized CIEDE2000. Font similarity applies each texture feature to a template glyph image and compares with FSIM. Background similarity reconstructs both backgrounds from spatial features and compares with MS-SSIM. The three scores are averaged into the final TAS.
Why This Matters
The paper targets a practical gap: content producers localizing ads, packaging, games, and films need to swap text in images while keeping the original look, and that need is not limited to English and Chinese. By releasing both the dataset and the metric, the work gives the field a way to train on real low-resource pairs and to evaluate style preservation when no reference image exists — a common situation in real deployments.
Real-world applications named or implied by the paper:
- Advertisement banners and product packaging localization
- Game and film localization
- Augmented reality signage and signage translation
- Reuse and adaptation of visual assets across linguistic and cultural contexts, including K-culture, Japanese pop content, and Arabic entertainment
Industry relevance: the two-stage recipe is cheap relative to the payoff — Stage 2 needs less than 5% of the synthetic data and 10% of the epochs, and runs in 0.3 hours on the reported hardware. That makes real-domain adaptation feasible for production pipelines rather than a research-only exercise. TAS also gives teams an interpretable signal about whether a failure came from color, font, or background, which image similarity metrics do not provide.
Future Directions
- Expand language coverage. The authors state that restricted language coverage may hinder generalization and that they plan to expand language diversity.
- Address long-text degradation. Editing performance declines for longer text inputs because of the scarcity of long-text samples in the collected data; related failure cases are analyzed in the Appendix.
- Collect more diverse real data. The limited size of real-world datasets and uncommon styles or noise conditions remain open problems.
- Explore unsupervised domain adaptation and zero-shot editing for unseen scripts to improve scalability and real-world applicability.
Target Audience
Researchers and engineers working on scene text editing, text image generation, diffusion-based image editing, and multilingual OCR-adjacent systems. It is also relevant to practitioners in localization, advertising, packaging, and AR who need to modify text in real images while preserving appearance, and to anyone looking for a dataset or an evaluation metric for low-resource text style preservation. Readers need some familiarity with diffusion models and OCR to follow the training and metric details.
Authors’ abstract
Scene Text Editing (STE) is the task of modifying text content in an image while preserving its visual style, such as font, color, and background. While recent diffusion-based approaches have shown improvements in visual quality, key limitations remain: lack of support for low-resource languages, domain gap between synthetic and real data, and the absence of appropriate metrics for evaluating text style preservation. To address these challenges, we propose STELLAR (Scene Text Editor for Low-resource LAnguages and Real-world data). STELLAR enables reliable multilingual editing through a language-adaptive glyph encoder and a multi-stage training strategy that first pre-trains on synthetic data and then fine-tunes on real images. We also construct a new dataset, STIPLAR(Scene Text Image Pairs of Low-resource lAnguages and Real-world data), for training and evaluation. Furthermore, we propose Text Appearance Similarity (TAS), a novel metric that assesses style preservation by independently measuring font, color, and background similarity, enabling robust evaluation even without ground truth. Experimental results demonstrate that STELLAR outperforms state-of-the-art models in visual consistency and recognition accuracy, achieving an average TAS improvement of 2.2% across languages over the baselines.