Skip to content
AI.info

Research

A Large-scale Dataset for Robust Complex Anime Scene Text Detection

Overview Research area: Computer vision - scene text detection datasets, specifically for anime imagery. Technical level: Intermediate. The paper is readable for anyone familiar with object detection

A Large-scale Dataset for Robust Complex Anime Scene Text Detection
arXiv
2510.07951
Published
2025-10-09
Authors
Ziyi Dong, Yurui Zhang, Changmao Li, Naomi Rue Golding, Qing Long

AI summary

Overview

Research area: Computer vision - scene text detection datasets, specifically for anime imagery.

Technical level: Intermediate. The paper is readable for anyone familiar with object detection basics, though the annotation-pipeline details (text-block definitions, CLIP fine-tuning) assume some background.

Scope: The paper introduces AnimeText, a large-scale anime scene text detection dataset of 735K images and 4.2M annotated text blocks with hierarchical annotations and hard negative samples, plus cross-dataset benchmarks on four detectors.

What This Paper Is About

Existing text detection datasets are built for natural or document scenes, where text uses regular fonts, orderly layouts, monotonous colors, and is usually aligned along straight or curved lines. Anime scenes break all of these assumptions: text is stylized or handwritten, scattered irregularly, multilingual, and frequently confused with decorative patterns and symbols that look like text. The paper's goal is to close this gap by releasing a large, high-quality anime text detection dataset and demonstrating that training on it substantially improves detection performance in anime scenes.

Key Contributions

  1. A large, diverse, multilingual anime scene text detection dataset. AnimeText contains 735K images and 4.2M annotated text instances across English, Chinese, Japanese, Korean, and Russian, described as 5x larger than existing text detection datasets, and focused specifically on text localization rather than recognition.
  2. Extensive cross-dataset benchmarking. Experiments show AnimeText works both as a training set that improves anime text detection across multiple baseline models and as a test set that presents a new challenge to the community.
  3. A scalable annotation workflow. A three-stage pipeline (manual box annotation, hard negative sample annotation, multi-granularity annotation) combining pre-trained models with human review, plus analysis of the dataset's statistics.
  4. Hard negative sample annotations. Explicit labels for text-like background elements (symbols, patterns, decorative marks), which are shown to improve precision when used for filtering and reweighting.

Main Findings

  • Large domain gap between natural and anime scenes. Detectors pre-trained on the natural-scene dataset ICDAR15, including Bridging Text Spotting, DBNet, and LRANet, degrade severely on anime text, with F1-scores dropping as low as 0.008 (DBNet, ICDAR15 train/test AnimeText).
  • Training on AnimeText bridges the gap. DBNet reaches an F1-score of 0.743 and LRANet reaches 0.855 when trained and tested on AnimeText. YOLO v11 trained and tested on AnimeText reaches Precision 0.878, Recall 0.825, F1 0.851, and mAP(50:95) 0.806.
  • The gap is bidirectional. YOLO v11 trained on AnimeText and tested on ICDAR15 scores Precision 0.187, Recall 0.310, F1 0.233, mAP(50:95) 0.083, confirming that anime-trained models also perform poorly on natural scenes.
  • Hard negative filtering improves precision. Applying the CLIP-H-based hard negative classifier to filter pseudo-labels led to a 26.9% increase in precision. In the ablation, YOLO v11 without hard negative filtering scored Precision 0.668, Recall 0.766, mAP 0.701; with filtering, 0.857 / 0.819 / 0.756; with hard negative sample reweighting, 0.878 / 0.825 / 0.806.
  • The classifier is highly accurate. The CLIP-H-based hard negative sample classifier, fine-tuned on 491 text patches and 473 hard negative patches and evaluated on 150 held-out samples, achieved 98.4% accuracy and a 98.1% F1-score.
  • Anime text sits at the periphery of images. The spatial heatmap analysis shows anime text is mostly located along the image periphery, whereas natural-scene text concentrates near the center, which may partly explain the poor performance of natural-scene detectors.
  • Anime images are more saturated and higher contrast. AnimeText has channel means of 0.677 (Red), 0.618 (Green), 0.612 (Blue) with standard deviations of 0.317, 0.319, 0.315, compared with ImageNet (0.485, 0.456, 0.406 / 0.229, 0.224, 0.225) and Total-Text (0.457, 0.427, 0.402 / 0.284, 0.277, 0.284).
  • Language distribution is skewed toward Japanese. Japanese accounts for 65.57% of instances, English 30.21%, Others 1.86%, Chinese 1.44%, Korean 0.62%, and Russian 0.30%.
  • Scale comparison. AnimeText is reported as 26.12x larger than TextOCR and 72.3x larger than ICDAR19-ArT in image count, and 4.69x and 84.71x larger in text instances, with an average of 5.77 text instances per image. The split is 514,144 train / 147,191 valid / 73,725 test images and 3,228,144 / 922,758 / 88,320 instances respectively.
  • Text density and resolution. Most images contain fewer than 10 text instances, but a substantial number contain between 50 and 100, and some more than 100. Most images have resolutions around 900².

Methodology in Plain English

The authors collected anime images containing text from multiple online sources under research-use licenses, then annotated them in three stages designed to keep manual effort manageable.

In Stage 1, roughly 50,000 images were manually boxed at the level of semantically coherent text blocks (a sentence-like cluster) rather than individual words. A text block is defined by a distance rule: characters within a block must be close relative to the block's average character width and height, with a threshold alpha of approximately 0.5. A preliminary YOLOv11 model, trained on a multilingual subset for 50 epochs with the Adam optimizer and a learning rate of 1e-3, generated pseudo-labels for the whole dataset at a 0.45 confidence threshold, which were then manually verified and corrected.

In Stage 2, the authors analyzed the model's mistakes, sorting them into exaggerated/stylized artistic text, low-contrast text, text-like patterns or symbols, and disordered character arrangements. Errors in the last two categories were labeled as hard negative samples. To scale this labeling, they fine-tuned a pretrained CLIP-H image encoder, updating only layers 26 through 32 and appending a classifier head whose weights were initialized from CLIP text-encoder embeddings with a [text, pattern] input, with the forward pass computed by cosine similarity. This classifier was then applied across the dataset.

In Stage 3, they built multi-granularity annotations: the finest level B⁰ is a text block, and coarser level B¹ groups closely arranged B⁰ boxes using an analogous distance rule with threshold beta. They also produced polygon annotations by running the Segment Anything Model (SAM) on existing box annotations and then manually refining the results.

For experiments, they ran cross-dataset evaluation with YOLO v11, DBNet, LRANet, and Bridging Text Spotting across ICDAR15 and AnimeText, and trained YOLO v11 for 50 epochs with batch size 512, the AdamW optimizer, learning rate 0.001, on 1 A100 GPU for 26 hours.

Why This Matters

Impact on research. The paper provides a benchmark for a domain that mainstream text detection datasets largely ignore. It deliberately decouples detection from recognition, so the annotations support localization research without requiring transcription, and it supplies hard negative samples that make text-versus-decoration discrimination an explicit, measurable task.

Real-world applications:

  • Anime and manga translation pipelines that need to locate text before recognizing and translating it.
  • Manga and anime restoration or remastering, where text regions must be identified so they can be edited or redrawn.
  • Multimodal large language model (MLLM) systems that need to ground textual content in images, including vision-language chat and visual question answering.
  • Accessibility tools such as assistive technologies for visually impaired users who need textual content in stylized imagery read aloud.

Industry relevance. The dataset was produced with participation from DeepGHS, an anime hobbyist syndicate, and is released on HuggingFace under a CC BY-NC-SA 4.0 license. Its scale and hierarchy annotations are directly relevant to media localization, content moderation, and any product that must process CJK-heavy stylized imagery at volume.

Future Directions

  • Extend to dynamic scenes. The current dataset covers only static anime scenes; adding temporal annotations for motion, motion blur, and dynamic compositions is left as future work.
  • Add transcriptions. AnimeText supports detection only. Adding text transcriptions would create a comprehensive anime OCR benchmark and enable end-to-end OCR and VQA tasks, which the authors identify as a current limitation.
  • Build universal detectors. The authors suggest investigating text detection datasets generalizable across both anime and natural scenes, since the measured domain gap runs in both directions.
  • Scale and refine hard negative annotations. The hard negative classifier was fine-tuned on 491 text patches and 473 hard negative patches; expanding and diversifying that supervision, and addressing the acknowledged variance from manual labeling thresholds, remain open questions.

Target Audience

Researchers and engineers working on scene text detection, OCR, and multimodal models who need a benchmark for stylized, multilingual, non-natural imagery. It is also useful for practitioners in anime/manga localization, digital restoration, and content tooling, and for dataset builders interested in semi-automated annotation pipelines that combine pre-trained models with human review.

Authors’ abstract

Current text detection datasets primarily target natural or document scenes, where text typically appear in regular font and shapes, monotonous colors, and orderly layouts. The text usually arranged along straight or curved lines. However, these characteristics differ significantly from anime scenes, where text is often diverse in style, irregularly arranged, and easily confused with complex visual elements such as symbols and decorative patterns. Text in anime scene also includes a large number of handwritten and stylized fonts. Motivated by this gap, we introduce AnimeText, a large-scale dataset containing 735K images and 4.2M annotated text blocks. It features hierarchical annotations and hard negative samples tailored for anime scenarios. %Cross-dataset evaluations using state-of-the-art methods demonstrate that models trained on AnimeText achieve superior performance in anime text detection tasks compared to existing datasets. To evaluate the robustness of AnimeText in complex anime scenes, we conducted cross-dataset benchmarking using state-of-the-art text detection methods. Experimental results demonstrate that models trained on AnimeText outperform those trained on existing datasets in anime scene text detection tasks. AnimeText on HuggingFace: https://huggingface.co/datasets/deepghs/AnimeText

Read the original paper