Skip to content
AI.info

Research

PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures

Overview Research area: Zero-Shot Anomaly Detection (ZSAD) in computer vision, specifically the use of vision-language models (CLIP) combined with prompt learning and Mixture-of-Experts (MoE) mechanis

PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures
arXiv
2511.18116
Published
2025-11-22
Authors
Yuheng Shao, Lizhang Wang, Changhao Li, Peixian Chen, Qinyuan Liu

AI summary

Overview

Research area: Zero-Shot Anomaly Detection (ZSAD) in computer vision, specifically the use of vision-language models (CLIP) combined with prompt learning and Mixture-of-Experts (MoE) mechanisms.

Technical level: Intermediate. The paper assumes familiarity with CLIP-style image-text alignment, prompt learning (static, learnable, and dynamic prompts), and the MoE paradigm of sparse expert routing.

Scope: The paper proposes PromptMoE, a framework that replaces monolithic text prompts with a visually-guided, sparsely routed mixture of expert prompts, and evaluates it across 15 industrial and medical anomaly detection datasets (arXiv:2511.18116v1 [cs.CV], 22 Nov 2025).

What This Paper Is About

Zero-Shot Anomaly Detection (ZSAD) asks a model to find and localize defects or abnormalities in images of object classes it never saw during training — for example, catching a scratch on a product category introduced after deployment. Recent methods adapt CLIP by learning "normal" and "abnormal" text prompts, but the authors argue these approaches hit a representational bottleneck: a single fixed prompt (or a single dynamic prompt generator) cannot capture the diversity of unseen anomalies, while simply adding more static prompts causes overfitting to the auxiliary training data. PromptMoE addresses this by learning a pool of expert prompts as composable semantic primitives and combining them per-image.

Key Contributions

  1. A compositional prompt-learning paradigm for ZSAD. PromptMoE replaces monolithic prompts with a Mixture-of-Experts mechanism that dynamically combines a learned basis set of semantic primitives, intended to improve generalization while mitigating overfitting.

  2. The Visually-Guided Mixture of Prompt (VGMoP) module. VGMoP uses an image-gated sparse router and separate expert prompt pools for normal and abnormal states to construct instance-specific textual prompts for each visual input.

  3. Two auxiliary regularization losses. An expert load-balancing loss (weighted by α) and an expert decoupling loss (weighted by β) are introduced to keep routing dynamic and to promote diversity among experts.

  4. Extensive evaluation across 15 datasets. Experiments span seven industrial benchmarks and eight medical datasets, reporting state-of-the-art average performance on both image-level and pixel-level metrics.

Main Findings

  • Industrial image-level results. PromptMoE reaches an average of 92.4 I-AUROC and 93.4 AP across the seven industrial datasets. On MVTec AD it scores 93.8 I-AUROC and 97.2 AP, described as +1.8% AUROC over the runner-up. Other reported values include VisA (85.0, 88.2), MPDD (82.3, 84.5), BTAD (93.4, 95.7), SDD (97.4, 94.0), DAGM (98.9, 96.4), and DTD-Synthetic (95.9, 98.1).

  • Industrial pixel-level results. The average is 96.2 pixel AUROC and 89.2 PRO. Individual values include MVTec AD (91.8, 83.2), VisA (95.6, 89.2), MPDD (96.8, 88.8), BTAD (94.9, 80.6), SDD (98.1, 95.6), DAGM (97.8, 93.9), and DTD-Synthetic (98.3, 93.2).

  • Medical image-level results. PromptMoE averages 97.4 AUROC and 97.5 AP across HeadCT, BrainMRI, and Br35H. On HeadCT specifically it reports 98.2 AUROC and 98.2 AP, which the authors state surpasses the next-best method by 3.4% in AUROC and 4.7% in AP — notable because the model was trained only on industrial data.

  • Medical pixel-level results. The average is 85.6 pixel AUROC and 68.8 PRO across ISIC (91.1, 81.4), CVC-ColonDB (84.3, 74.9), CVC-ClinicDB (84.5, 70.9), Kvasir (81.6, 48.8), and Endo (86.3, 68.0).

  • Composition beats enumeration. In the ablation on compositional prompting (Table 2, MVTec AD / VisA reported as I-AUC and PRO): a Static Prompt baseline gives 91.7 / 82.0 on MVTec AD and 82.4 / 88.0 on VisA; adding a Static Ensemble gives 92.2 / 82.9 and 83.3 / 88.3; swapping in VGMoP (single-layer, no auxiliary losses) gives 93.1 / 83.3 and 84.1 / 89.0; the full PromptMoE gives 93.8 / 83.2 and 85.0 / 89.2. The authors highlight the +0.9% I-AUC gain on MVTec AD from replacing the static ensemble with VGMoP as evidence that dynamic composition, not prompt count, drives the improvement.

  • Load balancing needs tuning. Removing the load-balancing loss entirely (α = 0) drops MVTec AD to 92.1 I-AUC / 82.5 PRO and VisA to 83.4 I-AUC / 88.2 PRO. α = 0.01 gives the best results (93.8 / 83.2 and 85.0 / 89.2), while an oversized α = 0.1 falls back to 92.9 / 83.1 and 84.2 / 89.2, which the authors attribute to over-penalizing router specialization.

  • Expert diversity and balancing are interdependent. Training curves (Fig. 7) indicate that without the decoupling loss, expert representations become redundant, the balancing loss degrades sharply, and I-AUC suffers. With both losses active, both auxiliary terms stay at healthy levels and final performance improves.

  • Normal and abnormal routing behave differently. The activation-frequency analysis shows the normal branch converging on a small, fixed set of core experts across all test datasets, while the abnormal branch stays dynamic and sparse. The authors note a general-purpose expert (E3) combined with dataset-specific experts — higher activation of E7 and E8 on VisA versus E1 and E6 on MPDD — and observe non-uniform gating weights rather than a simple average.

  • Multi-layer feature fusion helps. The layer ablation shows using all four visual encoder layers {6, 12, 18, 24} is most robust, with MVTec AD at 93.8 I-AUC / 83.2 PRO and VisA at 85.0 I-AUC / 89.2 PRO, versus single-layer and partial combinations that score lower.

Methodology in Plain English

The core idea. Instead of teaching a model one "normal" sentence and one "abnormal" sentence, PromptMoE builds a library of prompt fragments — the authors call them expert prompts or semantic primitives — and learns a small router network that picks the best few fragments for each specific image. A scratch on one product and a discoloration on another get different text descriptions, even though the total set of parameters stays small.

How the routing works. A set of learnable queries looks at the image's patch features through cross-attention, and the result is averaged into a single routing vector per layer. A small two-layer MLP turns that vector into scores over the expert pool; the top-k experts are selected and their outputs are combined using softmax-normalized weights. This is done independently at each selected layer of the frozen CLIP visual encoder, and independently for the normal and abnormal prompt pools.

How the prompts are assembled. The normal-state prompt consists of an aggregated normal state, a class token (the category name or a generic placeholder like "object"), and a set of learnable context tokens. The abnormal-state prompt is built by appending an aggregated abnormal state as a suffix to the normal state before the class and context tokens. The text encoder then produces embeddings that are compared with patch features to build the anomaly map.

How to keep the experts useful. Two extra losses are added during training. The load-balancing loss encourages all experts to contribute roughly equally across a batch, so the router does not collapse onto a few early winners. The decoupling loss pushes the mean representations of the experts in a layer toward orthogonality, so experts stay distinct rather than redundant. The total objective combines BCE for the image-level score, Dice and Focal losses for the pixel-level map, and these two auxiliary terms.

Experimental setup. CLIP ViT-L/14@336px is the frozen backbone; images are resized to 518×518; features come from layers {6, 12, 18, 24} of the 24-layer visual encoder. Training runs 15 epochs with Adam (betas 0.6, 0.999), batch size 16, learning rate 0.001, with warm-up over the first 3 epochs. VGMoP uses N_q = 8 queries, E = 8 experts per pool, top k = 4 selection, 8 attention heads, and a router hidden dimension of 256. State sequence lengths are M_n = 5 and M_a = 6, with context length M_q = 8. Auxiliary weights are α = 0.01 and β = 0.005; temperatures are τ = 0.07 (pixel-level) and τ' = 0.01 (image-level), and the final anomaly map is smoothed with a Gaussian filter with sigma 4. Everything runs on a single NVIDIA GeForce RTX 3090 with PyTorch 1.13.1.

Evaluation protocol. The model is primarily trained on MVTec AD and evaluated zero-shot on the other 14 datasets; for MVTec AD itself, it is trained on VisA, which has a disjoint set of object categories. Metrics are AUROC and AP at the image level, and pixel-wise AUROC and PRO at the pixel level. Baselines are WinCLIP, APRIL-GAN, AnomalyCLIP, AdaCLIP, and FAPrompt, with results cited from original papers or reproduced under a unified setting where values were missing.

Why This Matters

Impact on research. The paper reframes prompt engineering for ZSAD from "learn one better prompt" to "learn how to compose a prompt." If the compositional view holds up, it suggests a general recipe for adapting frozen vision-language models to open-ended concepts without enlarging the prompt budget — and the expert-activation analysis offers a concrete, interpretable picture of how a router distributes behavior between a stable concept (normality) and a variable one (anomaly types).

Real-world applications:

  • Industrial quality control: automated visual inspection of production lines for scratches, blemishes, and deformations, including newly introduced product categories that have no labeled defect examples.
  • Medical imaging triage: flagging and outlining anomalous regions in CT, MRI, and other modalities, where the model trained on industrial data still reports strong HeadCT and BrainMRI numbers.
  • Dermatology and endoscopy screening: pixel-level localization on datasets such as ISIC, CVC-ColonDB, CVC-ClinicDB, Kvasir, and Endo, where outlining the lesion boundary matters as much as detecting it.
  • Deployment-constrained inspection: because CLIP's parameters stay frozen and only the prompt pools and routers are trained, the approach avoids retraining a large backbone for each new product line.

Industry relevance. The author affiliations include Ant Group alongside Tongji University, and the paper emphasizes the practical scenario where new product categories and unforeseen defect types appear faster than labeled data can be collected — a recurring constraint in manufacturing and clinical settings.

Future Directions

  • What the truncated appendix would have shown. The provided content cuts off mid-sentence in the appendix's analysis of expert sharing strategies (referencing MMoE-style parameter sharing between the normal and abnormal branches). The full results of that investigation, plus the promised visualizations of expert semantics and category-level performance tables, are not reported in the available text.

  • Pushing toward more selective or larger expert pools. The paper fixes E = 8 and k = 4 per pool per layer without reporting a sweep over those values; whether the balance between capacity and overfitting shifts with larger pools is not reported.

  • Extending beyond the two-state assumption. The design assumes exactly one normal state and one abnormal state per instance. Whether the same compositional mechanism scales to multiple distinct anomaly families or finer-grained state taxonomies is an open question.

  • Reducing dependence on the auxiliary training set. The authors note that static prompts overfit auxiliary data, but the metrics for how much auxiliary data PromptMoE itself needs, and how it degrades with less of it, are not reported.

  • Broadening the modality coverage. All 15 datasets are 2D images. Whether the same prompt-mixture idea transfers to video, 3D, or multimodal industrial inspection streams is not explored.

Target Audience

This paper is most useful for computer vision researchers and graduate students working on anomaly detection, zero-shot or open-vocabulary recognition, and parameter-efficient adaptation of vision-language models. It also suits applied machine learning engineers in manufacturing and medical imaging who need models that generalize to new categories without retraining a backbone. Readers should already be comfortable with CLIP-style contrastive alignment, prompt learning, and the Mixture-of-Experts routing pattern to get the most out of the method section.

Authors’ abstract

Zero-Shot Anomaly Detection (ZSAD) aims to identify and localize anomalous regions in images of unseen object classes. While recent methods based on vision-language models like CLIP show promise, their performance is constrained by existing prompt engineering strategies. Current approaches, whether relying on single fixed, learnable, or dense dynamic prompts, suffer from a representational bottleneck and are prone to overfitting on auxiliary data, failing to generalize to the complexity and diversity of unseen anomalies. To overcome these limitations, we propose $\mathtt{PromptMoE}$. Our core insight is that robust ZSAD requires a compositional approach to prompt learning. Instead of learning monolithic prompts, $\mathtt{PromptMoE}$ learns a pool of expert prompts, which serve as a basis set of composable semantic primitives, and a visually-guided Mixture-of-Experts (MoE) mechanism to dynamically combine them for each instance. Our framework materializes this concept through a Visually-Guided Mixture of Prompt (VGMoP) that employs an image-gated sparse MoE to aggregate diverse normal and abnormal expert state prompts, generating semantically rich textual representations with strong generalization. Extensive experiments across 15 datasets in industrial and medical domains demonstrate the effectiveness and state-of-the-art performance of $\mathtt{PromptMoE}$.

Read the original paper