Research
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
Overview Research area: Computer Vision — streaming Video Anomaly Understanding (VAU), combining lightweight video anomaly detection with Multimodal Large Language Model (MLLM) reasoning. Technical le

- arXiv
- 2609.07941
- Published
- 2026-09-07
- Authors
- Chia-Hui Chen, Shih-Ying Yeh, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
AI summary
Overview
- Research area: Computer Vision — streaming Video Anomaly Understanding (VAU), combining lightweight video anomaly detection with Multimodal Large Language Model (MLLM) reasoning.
- Technical level: Advanced. The paper assumes familiarity with MLLMs, vision encoders (SigLIP-So400m), LoRA fine-tuning, KV-cache/memory compression, and streaming inference protocols.
- Scope: The paper proposes ReactVAU, a Slow-Fast decoupled framework that performs causal, real-time streaming video anomaly detection and explanation by gating a heavyweight 7B reasoning model behind a lightweight detector and an anomaly-aware memory.
What This Paper Is About
Existing top-tier Video Anomaly Understanding models rely on offline inference, observing the entire video before generating language, which breaks causality and makes them unusable on live surveillance streams. General streaming video models are causal but dilute rare, transient anomalies during memory compression and waste heavy MLLM computation on long normal intervals. ReactVAU addresses both problems by separating continuous lightweight anomaly filtering from on-demand heavyweight semantic reasoning, while protecting anomalous visual evidence from being compressed away.
Key Contributions
- ReactVAU framework: A Slow-Fast Decoupled Framework with a reactive event-gated mechanism that triggers heavyweight reasoning only upon detected anomalies, eliminating reliance on future frames and enabling real-time streaming VAU with reduced computational overhead.
- Spatial Grid Folding (SGF): A method that transforms short-term temporal anomaly detection into an efficient 2D spatial reasoning task, bypassing heavy temporal modeling by folding frames into grid images.
- Anomaly-Aware Persistent Memory (AAPM): A memory mechanism using an Anomaly Priority Score and a threat-based Anomaly Pool to protect transient abnormal features during continuous memory compression.
- Empirical validation: Experiments across multiple benchmarks showing competitive performance against state-of-the-art offline models under a strict causal streaming inference protocol, with significantly enhanced computational efficiency.
Main Findings
- Streaming detection performance: Under strict streaming constraints with no future information, ReactVAU reaches 88.44% AUC on UCF-Crime and 88.50% AP with 95.25% AUC on XD-Violence. It outperforms online methods such as MoniTor (82.57 AUC UCF-Crime; 55.01 AP / 79.11 AUC XD-Violence) and the fine-tuned StreamForest† baseline (85.26 AUC / 75.92 AP / 92.82 AUC).
- Competitiveness against offline models: ReactVAU remains competitive with the best offline fine-tuned methods, e.g. Holmes-VAD (89.51 UCF-Crime AUC; 90.67 XD-Violence AP), which can access future frames.
- Cost of forcing offline models online: When offline models are forced into online settings, performance degrades severely — Online-LAVAD drops to 76.06% AUC on UCF-Crime and 52.63% AP on XD-Violence.
- Long-range VAU gains on HIVAU-70K: ReactVAU† achieves CIDEr of 0.920 (Clip), 2.032 (Event) and 2.016 (Video), outperforming the fine-tuned StreamForest† (0.947 / 1.999 / 1.931) at Event and Video levels. Offline VADER† performs best at Clip level (CIDEr-C 1.040), and the paper attributes this to its global sampler observing the complete input and concentrating a fixed frame budget on anomaly-dense regions.
- Generic MLLMs degrade over long contexts: Zero-shot generic models collapse at Event and Video levels, e.g. InternVL2 scores 0.022 on CIDEr-E.
- Efficiency from decoupling: On the UCF-Crime test set (anomalies in approximately 43% of segments), heavyweight 7B LLM queries drop from 34,670 to 15,955, a 54.0% reduction.
- Latency: StreamForest runs at a constant 216.1 ms/query; ReactVAU runs at 22.1 ms/query during normal states and 269.5 ms/query when anomalies occur, giving an empirical weighted average of 98.3 ms/query on UCF-Crime.
- Real-world scalability projection: At a hypothetical 10% anomaly rate, queries fall to approximately 3,467 (90.0% reduction) with 39.9 ms/query; at 5%, approximately 1,734 queries (95.0% reduction) with 31.0 ms/query.
- SGF requires fine-tuning: Moving PaliGemma2-3B from single-frame inputs to zero-shot grid images disrupts pre-trained spatial priors, dropping XD-Violence AP from 53.71% to 43.42%; after fine-tuning on the Grid Image Dataset, AP surges to 79.55% (85.37 UCF-Crime AUC, 91.18 XD-Violence AUC).
- Ablation progression: StreamForest zero-shot reaches 78.09 AUC / 49.24 AP / 81.94 AUC; fine-tuning reaches 85.26 / 75.92 / 92.82; adding Slow-Fast decoupling with both modules tuned reaches 87.17 / 86.90 / 94.16; adding AAPM reaches 88.44 / 88.50 / 95.25.
- Score fusion matters: Purely replacing the fused score with the Slow score ("Replace") gives 87.28% AUC; adaptive fusion gives 88.08%; weighted 0.3D+0.7R gives 88.17%; the chosen 0.4D+0.6R gives 88.44%; 0.5D+0.5R gives 88.26%.
- Causal querying without resampling: ReactVAU answers a clip-level query at t = 20 s using 0–20 s memory and a video-level summary at t = 200 s using 0–200 s memory, without question-specific global resampling.
Methodology in Plain English
ReactVAU separates the job of watching a stream from the job of explaining what happened.
Fast Detection Module (SGF). A lightweight PaliGemma2-3B model continuously scans the stream at 4 FPS. Instead of using costly 3D temporal models, it takes a 1-second window of four frames (F1 to F4) and folds them into a single 2 × 2 Grid Image, letting the model's ordinary spatial attention do the temporal reasoning. The model was fine-tuned for 1 epoch on a Grid Image Dataset built from UCF-Crime and XD-Violence, updating only the projector and LoRA weights of the LLM. It outputs a detection score S_det computed as a localized softmax over the "Yes" and "No" target logits.
Anomaly-Aware Persistent Memory (AAPM). The framework builds on the StreamForest-7B memory hierarchy: Real-Time Perception (RTP) keeps a complete set of 729 tokens for the current frame; a Fine-grained Spatiotemporal Window (FSTW) stores 128 tokens per frame in a fixed 12-frame sliding window; and the Persistent Event Memory Forest (PEMF) stores 64 tokens per event node under a token quota of 2048, merging adjacent nodes with the lowest composite penalty. AAPM adds an Anomaly Priority Score, an exponential protective weight based on the detection scores of a node pair, to the total penalty so anomaly-rich events survive merging. It also adds an isolated Anomaly Pool of eight frames at 128 tokens per frame, held outside the 2048-token quota, which keeps the highest-scoring suspicious frames as evidence anchors. AAPM also adapts perception: it updates memory sparsely at 1 FPS during normal intervals and switches to dense encoding of all four frames once a trigger is passed.
Slow Reasoning Module. A dormant 7B MLLM is awakened only when S_det exceeds a trigger threshold. It retrieves a hierarchically ordered memory sequence combining PEMF, short-term FSTW, the Anomaly Pool and dense RTP tokens, and produces a secondary reasoning score S_reason plus a textual description. The final score is a weighted fusion, S_fused = 0.4·S_det + 0.6·S_reason, and the anomaly is confirmed only if S_fused exceeds a confirmation threshold, at which point a fine-grained description is generated autoregressively.
Why This Matters
Impact on research. The paper reframes VAU as a causal streaming problem rather than an offline video-analysis problem, and argues that accuracy–efficiency trade-offs should be evaluated on the Pareto frontier rather than accuracy alone. It also shows that generic streaming memory is poorly suited to anomaly domains because it dilutes rare transient features, motivating domain-specific memory protection.
Real-world applications:
- Live surveillance monitoring, where a system must flag and explain threats without waiting for the video to end.
- Traffic monitoring, where anomalies such as collisions or erratic driving are brief and rare relative to normal traffic.
- Public safety operations, where continuous streams from many cameras must be triaged under a fixed compute budget.
- Any deployment where compute is limited, since the trigger threshold can be tuned to trade detection sensitivity against latency.
Industry relevance. The efficiency profile is the central industrial argument: under a realistic 5% anomaly rate, the authors project a 95.0% reduction in 7B LLM queries and 31.0 ms/query average latency. The paper reports profiling on a single NVIDIA H100-80GB GPU and states that Supplementary Material profiles normal and triggered paths on an RTX 4090-24GB under a one-second streaming budget. The work is supported by the NVIDIA Taiwan AI Research & Development Center (TRDC), and authors are affiliated with National Tsing Hua University and NVIDIA Taipei.
Future Directions
- Closing the short-clip gap: ReactVAU ingests causally with one AAPM update per second, leaving short clips with less accumulated evidence than offline samplers; improving clip-level performance without breaking causality is an open problem the paper identifies.
- Threshold calibration and transfer: The trigger and confirmation thresholds are calibrated from the training split; the paper reports cross-dataset threshold transfer in Supplementary Material, but automatic or adaptive calibration remains a natural extension.
- Fusion strategy refinement: The paper shows that replacing the detection score with the reasoning score hurts AUC (87.28%), and that the best fixed weight is 0.4/0.6; more principled or learned fusion is suggested by the gap between weighted and adaptive strategies.
- Deployment beyond the tested hardware and rate assumptions: Latency at 5% and 10% anomaly rates is a theoretical projection rather than an empirical measurement, so validating ReactVAU on real multi-camera streams at those rates would test the scalability claims.
Target Audience
Researchers and engineers working on video anomaly detection, streaming video understanding, and multimodal LLM systems, particularly those interested in real-time surveillance deployment under fixed compute budgets. It is also relevant to practitioners designing gated or cascaded inference architectures, since the paper's main lesson — using a cheap model to gate an expensive one — transfers beyond the anomaly domain. Readers without background in MLLMs, memory compression or streaming inference will find the method sections dense.
Authors’ abstract
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). Existing VAU methods rely on offline inference with global temporal sampling, which violates causality and prevents deployment in live surveillance streams. Conversely, general streaming video models satisfy causal access but dilute rare transient anomalies during memory compression and often invoke heavyweight MLLMs uniformly over long normal intervals. React VAU addresses this gap with three synergistic components: a lightweight Fast Detection Module based on Spatial Grid Folding (SGF) for continuous anomaly filtering; an Anomaly-Aware Persistent Memory (AAPM) that protects critical visual cues from temporal decay; and a heavyweight Slow Reasoning Module that remains dormant during normal streams and is awakened only by suspicious events for semantic verification and causal description. Extensive experiments on multiple benchmarks demonstrate that ReactVAU operates under strict streaming constraints while simultaneously achieving competitive performance in both anomaly detection and causal reasoning, alongside significantly enhanced computational efficiency by minimizing heavyweight MLLM invocations. Project page is available at https://huiyuiui.github.io/ReactVAU/