Skip to content
AI.info

Research

AnoStyler: Text-Driven Localized Anomaly Generation via Lightweight Style Transfer

Overview Research area: Computer vision, specifically industrial visual anomaly detection and zero-shot anomaly image generation. Technical level: Intermediate. The paper combines familiar building bl

AnoStyler: Text-Driven Localized Anomaly Generation via Lightweight Style Transfer
arXiv
2511.06687
Published
2025-11-10
Authors
Yulim So, Seokho Kang

AI summary

Overview

Research area: Computer vision, specifically industrial visual anomaly detection and zero-shot anomaly image generation.

Technical level: Intermediate. The paper combines familiar building blocks (U-Net, CLIP, style transfer losses) in a novel arrangement, so readers with basic deep learning background can follow it, but the loss design and evaluation protocol assume some familiarity with anomaly detection benchmarks.

Scope in one sentence: The paper introduces AnoStyler, a zero-shot method that turns a single normal image into a realistic anomaly image by applying text-guided, mask-localized style transfer with a lightweight U-Net, and shows it improves downstream anomaly detection on MVTec-AD and VisA.

What This Paper Is About

Anomaly detection models need examples of defects to learn from, but real defect images are rare, diverse, and expensive to collect. Prior anomaly generation methods try to fill this gap, yet they generally suffer from at least one of three problems: generated anomalies look unrealistic, they require large amounts of real normal (or even anomalous) images, or they rely on memory-heavy diffusion models. AnoStyler's goal is to generate realistic, diverse, and semantically aligned anomalies from just one normal image plus a category label and defect type, using a much smaller model.

Key Contributions

  1. Style transfer reframing of anomaly generation. The paper frames zero-shot anomaly generation as text-guided localized style transfer, where a normal image is transformed by modifying visual attributes only in masked regions while preserving overall structure. The authors state that, to their knowledge, this is the first approach to effectively leverage style transfer for this purpose.

  2. A single-image, zero-shot framework. AnoStyler needs only one normal image (plus a category label and expected defect type) to synthesize each anomaly, removing the dependency on large collections of normal images that other methods require, and requiring no real anomaly images at all.

  3. A lightweight architecture and generation pipeline. The system combines a non-parametric, category-agnostic shape-guided mask generator, template-based two-class prompt generation, and a lightweight stylization network with tailored CLIP-based losses, achieving far lower parameter count and compute than diffusion-based baselines.

  4. Two new loss terms for spatial and semantic precision. The Mask-Weighted Co-Directional Loss (emphasizing masked regions in both global and patch-wise directional alignment) and the Masked CLIP Loss (aligning only the masked region of the output with the anomaly prompt set) are introduced to localize and semantically align the synthesized anomalies.

Main Findings

  • Anomaly generation quality on MVTec-AD: AnoStyler achieved the highest Inception Score (IS = 2.04) and the second-highest Intra-Cluster pairwise LPIPS distance (IC-L = 0.32) on MVTec-AD. The best IC-L on that dataset was AnomalyAny at 0.33.

  • Anomaly generation quality on VisA: AnoStyler obtained the best scores for both metrics on VisA (IS = 1.55, IC-L = 0.32), compared with competing zero-shot and few-shot baselines.

  • State-of-the-art zero-shot anomaly detection on MVTec-AD: AnoStyler reported I-AUC 98.0, I-AP 99.0, I-F1 97.0, P-AUC 94.4, P-AP 62.9, P-F1 60.7, and PRO 88.3. For comparison, the few-shot method AnoDiff reached I-AUC 99.2 and P-AUC 99.1 on the same dataset.

  • State-of-the-art zero-shot anomaly detection on VisA: AnoStyler reported I-AUC 93.9, I-AP 95.3, I-F1 90.1, P-AUC 93.8, P-AP 31.4, P-F1 36.4, and PRO 84.3. Its I-AUC of 93.9 exceeds every few-shot baseline listed (DFMGAN 83.7, AnoDiff 86.9, AnoGen 90.4).

  • Comparable to few-shot despite no real anomalies: The authors emphasize that AnoStyler, with no access to real anomaly images, yielded performance comparable with or even superior to few-shot baselines that do have access to a small number of real anomalies.

  • Ablation confirms all three loss components matter: On MVTec-AD, the baseline objective (a) gave IS 1.70, IC-L 0.25, I-AUC 88.2, P-AUC 85.7. Adding the modified global directional loss (b) raised these to 1.86, 0.29, 95.2, 92.5. Adding the patch-wise term as well (c) gave 1.96, 0.30, 96.7, 93.2. The full objective (d) achieved 2.04, 0.32, 98.0, 94.4. Qualitative examples show successive gains in realism and prompt alignment.

  • Large efficiency advantage: AnoStyler has 263M total parameters: 91M for SAM, 0.61M for the stylization network, 151M for the CLIP encoders, and 20M for the content-loss feature extractor. Diffusion-based baselines (AnoDiff, AnoGen, RealNet, AnomalyAny) each contain over 1B parameters.

  • Lower compute per generated anomaly: Generating one anomaly image takes 9.5 TFLOPs for AnoStyler versus 22.8 TFLOPs for AnomalyAny, the most directly comparable baseline.

  • Fast mask generation: Averaged over 100 independent runs, generating a single mask takes 0.54 ms (Line), 0.09 ms (Dot), and 115.23 ms (Freeform).

  • Qualitative behavior: AnoStyler's outputs are described as competitive with few-shot methods; heuristic zero-shot methods (CutPaste, DRAEM, NSA) tend to produce unrealistic anomalies, RealNet is reported to lack realism, and AnomalyAny produces more realistic anomalies but often introduces artifacts and overly smoothed details in some categories.

Methodology in Plain English

The pipeline takes one normal image, its category label (for example, "wood"), and a defect type (for example, "scratch"), and produces a synthetic defect image in three steps.

Step 1 — Build an anomaly mask procedurally. Instead of learning where defects go, the method draws random shapes from three "meta-shape priors": Line (straight or wavy strokes with variable thickness), Dot (radially sampled blobs that can be smooth, spiky, or irregular), and Freeform (unconstrained random trajectories resembling diffused or amorphous regions). The number of regions per image is drawn from a categorical distribution with exponentially decaying probabilities, controlled by a maximum region count and a decay coefficient, so fewer anomalies are favored but complex compositions occasionally occur. The primitive masks are unioned, then intersected with a foreground mask. For object-centric categories, the foreground mask comes from Segment Anything Model with a ViT-B backbone, using four corner points as positive prompts to extract the background and then negating it; for texture-centric categories, the entire image is treated as foreground.

Step 2 — Generate two sets of text prompts. Using templates, the method builds a normal prompt set and an anomaly prompt set from the category-defect pair. State descriptors such as "flawless [c]" and "[c] with [d] defect" are inserted into prompt templates like "a photo of the [s]", and embeddings of each set are averaged to reduce the bias of any single prompt. Default tokens ("sample" for category, "defect" for defect type) are used when information is missing.

Step 3 — Train a small stylization network. A lightweight U-Net (three downsampling blocks, three upsampling blocks) takes the normal image as input and is trained with CLIP-based losses while the CLIP text and image encoders stay frozen. The Mask-Weighted Co-Directional Loss compares the direction of change from normal to anomaly in image embedding space against the direction from normal to anomaly in text embedding space, both globally and in randomly cropped patches, with each patch weighted by how much of it the anomaly mask covers. The Masked CLIP Loss aligns only the masked region of the output with the anomaly prompt set. A content loss and a total variation loss keep the structure faithful and the output smooth. The final anomaly image is a composite: the network output inside the mask, and the untouched original image outside it.

Evaluation. Experiments ran on MVTec-AD (5,354 images, 10 object and 5 texture categories, 1 to 7 defect types each) and VisA (10,821 images, 12 object categories), with images and masks resized to 512 × 512. Generation quality was measured with IS and IC-L on 1,000 generated anomalies per category; downstream detection used a U-Net trained on 500 generated anomalies per category plus real normal images, evaluated with image- and pixel-level AUROC, AP, F1, and PRO. The stylization network used CLIP ViT-B/32 encoders, all experiments ran on a single NVIDIA RTX 2080Ti with 11 GB memory, and every experiment was repeated five times with different random seeds, reporting the average.

Why This Matters

Impact on research. The paper argues that anomaly generation need not be dominated by heavyweight diffusion pipelines, and demonstrates that a style-transfer formulation with a 0.61M-parameter generator can compete with models exceeding 1B parameters. It also contributes a reusable recipe for localized, text-conditioned image editing where the goal is to change only part of an image's appearance while preserving its content.

Real-world applications (grounded in the industrial inspection setting the benchmarks represent):

  • Manufacturing quality inspection, where production lines need detectors for defects that are too rare or too varied to collect in bulk.
  • Cold-start or novel-product deployment, where only a handful of good sample images exist and no defects have been observed yet.
  • Resource-constrained or edge inspection, where the paper's 11 GB single-GPU setup and 9.5 TFLOPs per image matter more than peak accuracy.
  • Flexible retargeting of inspection systems, since changing the category label and defect type in the text prompt changes what kind of anomaly is generated without collecting new data.

Industry relevance. Quality control is a major cost center in manufacturing, and the paper's framing — a compact, single-image, prompt-driven generator — targets deployment scenarios where diffusion models are impractical due to memory or latency. The authors note the method produces a synthetic anomaly one image at a time from one normal image, which suits pipeline-style generation rather than large offline training runs.

Future Directions

  1. Defect type granularity in domains without labels. On VisA, annotations contain multiple free-form textual descriptions per anomaly image, so the authors uniformly assign the label "defect" to all anomaly images across all categories. Extending AnoStyler to exploit rich, free-form defect descriptions rather than a single generic token is a natural next step.

  2. Beyond the three meta-shape priors. Line, Dot, and Freeform are chosen to span plausible anomaly geometries, but the paper does not report how well they cover defect morphologies in domains outside MVTec-AD and VisA. Whether additional primitives, or learned shape priors, would improve generality is an open question.

  3. Scalability to settings beyond industrial inspection. Both benchmark datasets are industrial, and the evaluation protocol generates 1,000 and 500 anomalies per category for the two evaluation tasks. Whether the same pipeline holds for medical imaging, remote sensing, or other anomaly domains is not reported.

  4. Prompt sensitivity and category dependence. Prompt averaging is used to reduce the bias of individual prompts, and the paper notes default tokens when labels are missing, but it does not report a systematic study of how much performance depends on prompt wording. Per-category results are placed in Appendices E and F and are not included in the provided text.

Target Audience

This paper is most useful to computer vision researchers working on anomaly detection and anomaly synthesis, particularly those interested in zero-shot or few-shot generative augmentation for industrial inspection. It also suits practitioners building defect-detection systems who are constrained by GPU memory or latency and cannot deploy diffusion-based generators, and researchers studying text-driven localized image editing who want a compact application case of CLIP-guided style transfer. Readers looking for a full derivation of the mask-generation algorithms should note that Appendix B is only partially present in the supplied content, and the detailed per-category results in Appendices E and F are not included.

Authors’ abstract

Anomaly generation has been widely explored to address the scarcity of anomaly images in real-world data. However, existing methods typically suffer from at least one of the following limitations, hindering their practical deployment: (1) lack of visual realism in generated anomalies; (2) dependence on large amounts of real images; and (3) use of memory-intensive, heavyweight model architectures. To overcome these limitations, we propose AnoStyler, a lightweight yet effective method that frames zero-shot anomaly generation as text-guided style transfer. Given a single normal image along with its category label and expected defect type, an anomaly mask indicating the localized anomaly regions and two-class text prompts representing the normal and anomaly states are generated using generalizable category-agnostic procedures. A lightweight U-Net model trained with CLIP-based loss functions is used to stylize the normal image into a visually realistic anomaly image, where anomalies are localized by the anomaly mask and semantically aligned with the text prompts. Extensive experiments on the MVTec-AD and VisA datasets show that AnoStyler outperforms existing anomaly generation methods in generating high-quality and diverse anomaly images. Furthermore, using these generated anomalies helps enhance anomaly detection performance.

Read the original paper