Skip to content
AI.info

Research

PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates

PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates Overview Research area: Computer vision, specifically video anomaly detection (VAD) — the task of identify

PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
arXiv
2512.06845
Published
2025-12-07
Authors
Satoshi Hashimoto, Yanan Wang, Hitoshi Nishimura, Mori Kurokawa

AI summary

PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates

Overview

Research area: Computer vision, specifically video anomaly detection (VAD) — the task of identifying abnormal events such as assaults, traffic accidents, and theft in surveillance video.

Technical level: Advanced. The paper combines image-to-video diffusion models, vision–language models (CLIP, Qwen), Multiple Instance Learning (MIL), domain-adversarial training, and memory-bank architectures.

Scope: The paper proposes PA-VAD, a framework that trains a video anomaly detector using only real normal videos plus diffusion-synthesized pseudo-abnormal videos, eliminating any need for real abnormal footage during training.

What This Paper Is About

Video anomaly detection systems are usually trained either on normal data alone (unsupervised VAD) or on large collections of real abnormal videos with video-level labels (weakly supervised VAD). Both paths are limited: unsupervised methods can be brittle, and real abnormal footage is rare, costly to collect, and often constrained by safety and privacy rules.

PA-VAD's goal is to remove the real-abnormal-data requirement entirely. It generates class-aware pseudo-abnormal videos from a small set of real normal images, then trains a detector on real normal videos plus those synthesized clips — pairing this with a regularization module designed to fix a bias the authors discovered in the synthesized data.

Key Contributions

  1. A generation-driven VAD framework trained without any real abnormal videos. PA-VAD trains on real normal videos plus diffusion-generated pseudo-abnormal videos, while evaluation follows standard benchmark splits and protocols with real normal and abnormal videos.

  2. The Class-Aware Pseudo-Anomaly Generator (CA-PAG). A video-diffusion-based pseudo-anomaly generator that uses CLIP-guided initial image selection and VLM-driven prompt refinement to produce high-fidelity pseudo-abnormal clips anchored to class-relevant scenes.

  3. The Domain-Aligned Regularized Module (DARM). An adaptive memory module that mitigates a "pseudo-induced large-magnitude bias" through domain alignment (a DANN-based objective) and usage-aware memory updates, enabling stable MIL training with synthesized anomalies.

  4. Extensive benchmark evidence. Experiments on ShanghaiTech, UCF-Crime, and XD-Violence show state-of-the-art results against UVAD methods and performance that surpasses strong methods depending on real abnormal videos.

Main Findings

  • Headline detection results: PA-VAD achieves 98.2% AUC on ShanghaiTech, 82.5% on UCF-Crime, and 95.1% on XD-Violence under the Real/Pseudo setting (no real abnormal training videos).

  • Beats real-abnormal baseline methods on two benchmarks: PA-VAD surpasses the best real-abnormal WVAD baselines on ShanghaiTech and XD-Violence by +0.6% and +0.9%, respectively. On ShanghaiTech it exceeds CMRL (97.6%) by +0.6 points; on XD-Violence it exceeds UR-DMU (94.2%) by +0.9 points.

  • Improves on UVAD state of the art on UCF-Crime: PA-VAD surpasses MGSTRL (80.6%) by +1.9 points on UCF-Crime, even though its Real/Pseudo setting is described as naturally disadvantaged against Real/Real baselines there.

  • A discovered spatiotemporal magnitude bias: In frozen backbone features on UCF-Crime, real anomalies show mean ℓ2 norms of 20.52 (I3D), 24.22 (Qwen), and 19.14 (C3D), versus pseudo anomalies at 23.03, 25.55, and 20.41 respectively. This mild gap is strongly amplified where MIL Top-k operates.

  • DARM corrects the amplification: In the detector space on SHT, pseudo-abnormal norms inflate to 199.48 versus 9.13 for real abnormal without DARM — over 20 times larger. With DARM, pseudo-abnormal norms fall to 8.39 (SHT), 12.80 (Crime, vs. 11.06 real), and 33.63 (XD, vs. 39.08 real). Domain alignment alone only partially narrows the gap (SHT: 183.50).

  • DARM consistently improves the backbone: Replacing plain UR-DMU with DARM raises AUC from 96.0% → 98.2% on SHT, 80.2% → 82.5% on Crime, and 92.5% → 95.1% on XD, with a consistent gain also holding for the Sultani et al. classifier (93.7% / 77.8% / 81.8%).

  • Prompt refinement improves generation quality: On ShanghaiTech, refinement reduces FVD from 701 to 604 (a 97-point reduction) and KVD from 57.4 to 34.7 (22.7 points), while FID improves from 83.9 to 78.0 and KID stays unchanged at 0.03 — consistent with per-frame appearance statistics being identical while motion coherence improves.

  • Ablation on SHT: Random-init baseline 86.7% → CLIP-based initial image selection 94.9% (+8.2 points) → adding prompt refinement 96.0% → usage-aware memory update boosts to 97.6–97.7% → combining domain alignment and usage-aware update gives the best 98.2%. Domain alignment alone provides only modest improvement; the usage-aware update is the primary driver.

  • Robustness to the number of synthesized clips (SHT): Performance rises steadily from 85.4% at 14 clips, 94.6% at 35, 96.9% at 70, 97.3% at 105, to 98.2% at 140 clips, then slightly drops to 97.7% at 175 clips, suggesting mild saturation.

  • Open-set generalization to unseen anomaly classes: Under OpenVAD's open-set protocol, PA-VAD outperforms OpenVAD across all settings. On XD with only one seen class: 89.08 ± 4.52 versus OpenVAD's 72.50 (+16.6 points); with four seen classes: 94.99 ± 0.13 versus 88.25. On Crime: 77.65 ± 3.29 (1 seen), 80.86 ± 2.03 (3), 82.18 ± 0.25 (6), 82.30 ± 0.14 (9), versus OpenVAD's 76.73 / 77.78 / 78.82 / 80.14 and RTFM's 75.91 / 76.98 / 77.68 / 79.55.

  • Comparison with single-frame supervision: PA-VAD remains competitive with SF-VAD/FPL on SHT (98.2 vs. 98.4) despite not using the precise abnormal-frame annotations that SF-VAD relies on.

  • Stated limitation: On UCF-Crime, PA-VAD does not surpass the state of the art that trains with real abnormal videos under the WVAD setup. The authors attribute this to the difficulty of synthesizing minute-long, context-heavy anomalies with extended temporal dependencies, which are more prevalent in UCF-Crime than in ShanghaiTech.

Methodology in Plain English

The core pipeline. Training data consists only of real normal videos plus the names of abnormal classes that might appear at inference time. No real abnormal videos are used. The authors generate fake "abnormal" videos with a video diffusion model, then train a standard weakly supervised detector on real normal plus synthesized videos, evaluating on real normal and real abnormal test footage.

Step 1 — Pick good starting images. Randomly chosen normal frames can be semantically wrong for a target class (for example, an indoor home scene used to generate a "Road Accident"). CA-PAG instead scores normal images with CLIP using a positive text query (the class name plus surveillance-style phrases) minus a weighted negative query (suppressing nuisances like black screens or channel logos), then takes the Top-K with a scene-balanced ranking scheme so that populous camera/location scenes do not dominate selection.

Step 2 — Rewrite the prompt with a vision–language model. A raw class name such as "shoplifting" is ambiguous. A VLM looks at the selected initial image and produces a short, class-consistent abnormal description (for example, describing an individual removing merchandise, concealing it in a backpack, and evading cameras), which is concatenated with template prompts and fed to the video diffusion model along with the initial image.

Step 3 — Detect and fix the bias. When synthesized clips are plugged into a standard weakly supervised MIL pipeline, the fake anomalies tend to have exaggerated motion and viewpoint changes, inflating feature norms and skewing the MIL Top-k selection toward a few high-magnitude pseudo instances. DARM counters this with two mechanisms: (i) a domain-alignment objective that adversarially aligns the real-normal and pseudo-normal feature distributions, and (ii) a usage-aware memory update that measures how much each abnormal memory slot is used and pulls under-used slots more strongly toward their responsibility-weighted centers, so a few high-norm slots cannot monopolize Top-k selection. DARM is added on top of the UR-DMU backbone, which supplies a GL-MHSA-based encoder, dual normal/abnormal memory banks, and a MIL scoring head with uncertainty control.

Implementation choices reported: the Wan 2.2 image-to-video diffusion model (I2V-A14B) at 832 × 480 resolution with frame length 81; Qwen3 30B-A3B for prompt refinement; Qwen2.5-VL 7B-Instruct's final projector embedding (3,584-D) as an auxiliary feature extractor; and CLIP ViT-B/32 for initial image selection. Pseudo-video synthesis is performed once offline; generating 60 clips takes about 10 hours on two RTX 6000 Ada GPUs (96GB).

Datasets and synthesized volumes: ShanghaiTech (13 locations, 437 clips, 14 abnormal categories, 63 abnormal clips in the standard training split) uses 10 pseudo-abnormal clips per class (total 140) plus 50 pseudo-normal clips for domain alignment. UCF-Crime (about 1,900 videos, roughly 128 hours, 13 abnormal classes, 810 abnormal clips in its training split) uses 20 pseudo clips per class (total 260) plus 80 pseudo-normal clips, with 10-crop augmentation. XD-Violence (4,754 untrimmed videos, 217 hours, 6 violence categories, official split of 3,954 training and 800 testing videos) uses 10 pseudo-abnormal clips for each of 6 classes plus 50 pseudo-normal clips.

Why This Matters

Impact on research. The paper reframes the data-scarcity problem in VAD: instead of reducing reliance on real abnormal footage, it removes that reliance entirely while still training under standard weakly supervised evaluation splits. It also contributes a concrete, measurable diagnosis — the spatiotemporal magnitude bias in synthesized anomalies and its amplification inside MIL Top-k optimization — that is relevant to anyone training detectors on generated data, not just VAD researchers.

Real-world applications:

  • Public safety and surveillance camera networks, where abnormal events such as assaults, explosions, and burglaries are exactly the events that are rare, dangerous, and hard to archive at scale.
  • Privacy-constrained settings such as hospitals, homes, or private premises, where collecting and storing footage of real incidents raises legal and ethical barriers.
  • Retail loss prevention, where classes like shoplifting and theft can be synthesized and adapted to specific store layouts without waiting for real incidents.
  • Traffic and roadway monitoring, where classes such as road accidents and vehicle collisions are explicitly used as generation targets in the paper's examples.

Industry relevance. The work comes from KDDI Research, Inc., a telecom research lab, and targets a practical deployment bottleneck: the cost, rarity, and privacy burden of abnormal footage. The reported synthesis cost (about 10 hours on two RTX 6000 Ada GPUs for 60 clips) is a one-time offline expense, and the paper shows that competitive performance is reached with only 70 synthesized clips on ShanghaiTech (96.9% AUC) — around the scale of the 63 real abnormal clips in the standard split.

Future Directions

  • Improving motion priors for pseudo generation. The authors explicitly name this as future work, motivated by the UCF-Crime limitation where minute-long, context-heavy anomalies with extended temporal dependencies remain hard for current video generators.
  • Extending evaluation to broader video generative models and settings. The current study uses one image-to-video generator (Wan 2.2 I2V-A14B) and one prompt-refinement VLM (Qwen3 30B-A3B); generalization across generators is untested in the reported content.
  • Closing the remaining UCF-Crime gap against Real/Real WVAD methods, which the paper reports as the one benchmark where PA-VAD does not surpass state-of-the-art methods that train with real abnormal videos.
  • Pushing open-set transfer further. DARM's usage-aware updates are credited with anti-overfitting behavior against seen-class specialization; the paper's open-set results are measured under a small number of random class splits (three), so broader and more systematic unseen-anomaly evaluation is a natural next step.

Target Audience

Researchers and graduate students in computer vision and video understanding, particularly those working on anomaly detection, weakly supervised learning, or synthetic-data training. It is also relevant to practitioners building surveillance, safety, or security analytics systems who face constraints on obtaining real abnormal footage, and to researchers studying domain adaptation, memory-augmented MIL architectures, or the reliability of diffusion-generated training data. A reader needs prior familiarity with MIL formulations, diffusion-based video generation, and vision–language models to follow the technical details, since the paper assumes that background rather than introducing it.

Authors’ abstract

Deploying video anomaly detection (VAD) in the real world is often constrained by the scarcity, privacy, and cost of collecting real abnormal footage. We propose PA-VAD, a novel pseudo-only framework that trains an anomaly detector without using any real abnormal videos, by pairing real normal videos with diffusion-synthesized pseudo-abnormal videos generated from a small set of real normal images. Beyond proposing a generation-driven training pipeline, we make a key empirical discovery: pseudo anomalies exhibit a characteristic spatiotemporal magnitude bias in feature space, which can dominate Multiple Instance Learning and degrade generalization if left unaddressed. To counter this pseudo-induced bias, we introduce the Domain-Aligned Regularized Module (DARM), which combines domain alignment with usage-aware memory updates to balance prototype coverage and stabilize optimization under biased pseudo supervision. Extensive experiments demonstrate that PA-VAD achieves 98.2% AUC on ShanghaiTech, 82.5% on UCF-Crime, and 95.1% on XD-Violence, and further improves generalization to unseen anomaly classes in open-set evaluations. Notably, PA-VAD surpasses the best real-abnormal WVAD baselines on ShanghaiTech and XD-Violence by +0.6% and +0.9%, respectively, and improves over the UVAD state of the art on UCF-Crime by +1.9% -showing that high-accuracy VAD is attainable without collecting real abnormal videos.

Read the original paper