Skip to content
AI.info

Research

AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

Overview Research area: Computer vision, specifically zero-shot visual anomaly detection and parameter-efficient adaptation of vision foundation models. Technical level: Advanced. The paper assumes fa

AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
arXiv
2601.20524
Published
2026-01-28
Authors
Matic Fučka, Vitjan Zavrtanik, Danijel Skočaj

AI summary

Overview

  • Research area: Computer vision, specifically zero-shot visual anomaly detection and parameter-efficient adaptation of vision foundation models.
  • Technical level: Advanced. The paper assumes familiarity with transformer backbones, LoRA-style adapters, flow-matching/diffusion image generation, and AUROC/F1 evaluation protocols.
  • Scope: The paper introduces AnomalyVFM, a model-agnostic framework that converts pretrained vision foundation models (DINOv2, DINOv3, RADIO) into zero-shot anomaly detectors via a synthetic data generator and parameter-efficient fine-tuning.

What This Paper Is About

Zero-shot anomaly detection means finding and localising defects in images of object classes the model has never seen during training. Existing state-of-the-art methods lean on vision–language models such as CLIP, while pure vision foundation models (VFMs) like DINOv2 have lagged behind. The paper argues this gap comes from two practical issues — insufficient diversity in auxiliary anomaly datasets and overly shallow VFM adaptation — and proposes a framework that fixes both, with RADIO as the backbone reaching an average image-level AUROC of 94.1% across 9 diverse datasets, surpassing prior methods by 3.3 percentage points.

Key Contributions

  1. AnomalyVFM framework: A model-agnostic method that transforms any pretrained VFM with a transformer backbone into a competitive zero-shot anomaly detector, using parameter-efficient adaptation rather than only training a final head.
  2. Three-stage synthetic dataset generator: Uses modern generative models (FLUX) to (i) create diverse anomaly-free object images, (ii) synthesise realistic local defects by inpainting at sampled locations, and (iii) filter samples with a feature-based verification step — requiring no normal or abnormal real samples.
  3. Parameter-efficient adaptation with a confidence-weighted loss: Low-rank (LoRA) adapters injected into query, value, and output projection layers throughout the backbone, a lightweight convolutional decoder producing pixel-level anomaly and confidence maps, and a loss that downweights ambiguous supervision.
  4. Broad validation: Generalisation demonstrated across three VFMs and 9 industrial benchmarks, plus transfer to 9 medical datasets without medical fine-tuning and competitive few-shot performance after minimal additional fine-tuning.

Main Findings

  • State-of-the-art zero-shot detection: AnomalyVFM with RADIO reaches an average image-level AUROC of 94.1% and image-level F1-Max of 87.6% across 9 industrial datasets, exceeding the next-best method (Bayes-PFL, 90.8% image-level AUROC) by 3.3 percentage points.
  • Localisation gains: AnomalyVFM attains an average pixel-level AUROC of 96.9%, improving prior methods by 0.9 percentage points, though its average pixel-level F1-Max of 44.3% trails AdaCLIP's 51.0%.
  • Generalises across VFMs: Under the adaptation strategy plus synthetic data, DINOv2 improves from 83.0% to 90.2% image-level AUROC, DINOv3 from 85.3% to 91.5%, and RADIO from 89.1% to 94.1%. Averaged over the VFMs, image-level AUROC rises by 6.1 p.p. and pixel-level AUROC by 10.7 p.p.
  • Synthetic data alone suffices: AnomalyVFM is trained solely on automatically generated data (10,000 images; 100 object types), whereas compared methods are trained on auxiliary real datasets such as the VisA or MVTec AD test sets.
  • Transfers to medical imaging: Without medical fine-tuning, AnomalyVFM reaches an average image-level AUROC of 94.0% and pixel-level AUROC of 90.0% on 9 medical datasets, improving pixel-level AUROC over prior methods by 1.2 p.p. Note AdaCLIP was also trained with auxiliary medical datasets.
  • Strong few-shot behaviour: After 50 additional fine-tuning iterations on a few normal samples, AnomalyVFM achieves image-level AUROC of 97.2 / 97.9 / 98.2 on MVTec AD for 1/2/4 shots, and 93.8 / 94.2 / 94.5 on VisA, matching or surpassing recent few-shot methods such as INP-Former.
  • Filtering is critical: Removing the dataset filtering step costs 3.8 p.p. image-level AUROC and 14.6 p.p. pixel-level AUROC; removing foreground-restricted anomaly placement costs 1.4 p.p. and 5.8 p.p. respectively.
  • Robust component choices: Replacing FLUX with QWEN-Image or WAN changes image-level AUROC by only −0.1 and −0.4 p.p.; swapping LoRA for AdaLN or VPT costs 0.7 and 1.0 p.p. respectively. Removing the confidence loss costs 0.6 p.p. image-level and 2.0 p.p. pixel-level AUROC.
  • Efficient inference: AnomalyVFM runs at 20.5 ms per sample on an NVIDIA A100, versus 82.4 ms for AdaCLIP and 208.5 ms for Bayes-PFL. The model has 345.8 million parameters, of which 35.4 million are trainable.
  • Training cost profile: Model training takes about two hours on a single A100 GPU, while synthetic dataset generation takes approximately one day on an A100 — the stated main bottleneck.

Methodology in Plain English

The pipeline has two halves. The first half builds training data from nothing: a text-to-image model (FLUX) generates close-up photos of objects on backgrounds, using object and texture tags drawn from lists of 100 objects and 50 backgrounds produced by an LLM (GPT-4o). A salient object segmentation network (IS-Net) extracts the object, a rectangle is sampled on the foreground, and the generator is re-prompted with an anomalous description (e.g. cracked, damaged, smudged, rotten) to inpaint a defect inside that region using the RePaint approach. Because generators do not always follow prompts, DINOv2 features of the clean and anomalous images are compared by cosine distance; if the maximum distance exceeds a threshold (T = 0.3), the sample is kept, and the thresholded distance map becomes the anomaly mask.

The second half trains the detector. The image passes through a pretrained VFM backbone whose transformer blocks are augmented with LoRA adapters (rank 64) in the query, value, and output projection layers. Features from the final block go into a small convolutional decoder with two upsampling blocks (convolution, GroupNorm, ReLU, bilinear upsampling); a final convolution produces both a segmentation map and a confidence map, while the backbone's [CLS] tokens feed a linear layer that predicts an image-level anomaly score. Training combines a Focal loss for the image score with a segmentation loss of L1 plus Focal (β = 5), weighted by the predicted confidence (α = 0.1) so that noisy synthetic labels contribute less. The whole model trains for 500 iterations with batch size 32 using AdamW at a learning rate of 10⁻⁴, with RADIOv2.5 ViT-L (patch size 16) as the default backbone and inputs resized to 768×768.

Why This Matters

  • Research impact: The paper challenges the assumption that high-level vision–language concept knowledge is necessary for zero-shot anomaly detection, showing that purely visual foundation models can lead when paired with diverse training data and deeper adaptation. It also provides a dataset generation recipe that competing methods can reuse — the supplementary material reports results for AACLIP, AnomalyCLIP, FAPrompt, AdaCLIP, and Bayes-PFL trained on the proposed synthetic dataset.
  • Real-world applications:
    • Manufacturing quality control, where new product lines appear without labelled defect examples (MVTec AD, VisA, BTAD, MPDD, Real-IAD, KSDD, KSDD2, DAGM benchmarks).
    • Texture and surface inspection (DTD-Synthetic) for materials like fabric, tile, and wood.
    • Medical imaging triage and lesion localisation (HeadCT, BrainMRI, BR35H, ISIC, ClinicDB, ColonDB, Kvasir, Endo, TN3K), achieved without any medical training data.
    • Road obstacle and infrastructure inspection, a domain cited in the paper's framing of anomaly detection.
  • Industry relevance: Inference at 20.5 ms per sample on an A100 and roughly two hours of training make deployment plausible for industrial inspection lines; the one-day dataset generation cost is a one-time investment reusable across backbones, and the authors report good performance is achievable with fewer than 10,000 generated images.

Future Directions

  • Improving the realism of generated defects and the accuracy of the resulting anomaly masks, since some anomaly-free images still pass the filter and could be removed by using a trained AnomalyVFM itself as a data filter.
  • Expanding and refining the LLM-generated [Object] tag lists, which the authors hypothesise would further improve performance.
  • Adapting the generator for medical domains — current pretrained generators failed to produce realistic medical images suitable for zero-shot training, so fine-tuning on an auxiliary medical imaging dataset is suggested.
  • Extending to RGB-D anomaly detection by integrating depth from monodepth foundational models such as Marigold, and using AnomalyVFM as a backbone for future few-shot and full-shot models.

Target Audience

Researchers and engineers working on industrial visual inspection, medical image analysis, and zero-shot or few-shot anomaly detection; practitioners who want to adapt pretrained vision foundation models with limited labelled data and modest compute; and readers interested in synthetic data generation pipelines for training data-scarce vision tasks.

Authors’ abstract

Zero-shot anomaly detection aims to detect and localise abnormal regions in the image without access to any in-domain training images. While recent approaches leverage vision-language models (VLMs), such as CLIP, to transfer high-level concept knowledge, methods based on purely vision foundation models (VFMs), like DINOv2, have lagged behind in performance. We argue that this gap stems from two practical issues: (i) limited diversity in existing auxiliary anomaly detection datasets and (ii) overly shallow VFM adaptation strategies. To address both challenges, we propose AnomalyVFM, a general and effective framework that turns any pretrained VFM into a strong zero-shot anomaly detector. Our approach combines a robust three-stage synthetic dataset generation scheme with a parameter-efficient adaptation mechanism, utilising low-rank feature adapters and a confidence-weighted pixel loss. Together, these components enable modern VFMs to substantially outperform current state-of-the-art methods. More specifically, with RADIO as a backbone, AnomalyVFM achieves an average image-level AUROC of 94.1% across 9 diverse datasets, surpassing previous methods by significant 3.3 percentage points. Project Page: https://maticfuc.github.io/anomaly_vfm/

Read the original paper