Skip to content
AI.info

Research

Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding

Overview Research area: Computer vision for traffic-safety video analysis, specifically zero-shot accident understanding in surveillance footage using vision-language models (VLMs) and open-vocabulary

Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding
arXiv
2606.12047
Published
2026-06-10
Authors
Tarandeep Singh, Soumyanetra Pal, Soham Biswas, Nishanth Chandran

AI summary

Overview

Research area: Computer vision for traffic-safety video analysis, specifically zero-shot accident understanding in surveillance footage using vision-language models (VLMs) and open-vocabulary detection.

Technical level: Advanced. The paper assumes familiarity with contrastive vision-language encoders (CLIP-style retrieval), generative VLMs, open-vocabulary object detection, and metric aggregation schemes.

Scope: The paper describes a three-stage zero-shot pipeline that predicts when an accident happens, what type it is, and where the impact occurs in the frame, and evaluates it on the zero-shot ACCIDENT@CVPR 2026 benchmark.

What This Paper Is About

In surveillance and fleet-safety video, knowing that an accident occurred is not enough: systems must also say when it happened, what category of collision it was (rear-end, T-bone, head-on, sideswipe, or single-vehicle), and where in the image the impact took place. The authors argue that asking a single vision-language model to answer all three questions in one prompt produces unstable predictions, so they split the problem into a when / what / where pipeline that never requires labelled real-world training data. The goal is reliable zero-shot generalization to unconstrained CCTV footage that is low resolution, compressed, occluded, and shot from shallow angles.

Key Contributions

  1. A three-stage when / what / where framework for zero-shot accident understanding from CCTV video, in which each stage can be upgraded independently of the others.
  2. A five-prompt classification scheme (baseline, motion, geometry, contrast, tiebreaker) with an entropy-gated pairwise adjudicator that resolves ambiguous votes without any labelled calibration data.
  3. A type- and scene-conditioned spatial localization strategy using an open-vocabulary detector, with score-weighted centroid aggregation across keyframes.
  4. A demonstration that this decomposition improves the harmonic-mean ACCIDENT score over centre-of-frame and Molmo2 pointing baselines, and evidence of a foreground bias in existing pointing models.

Main Findings

  • Overall benchmark result: On the ACCIDENT@CVPR 2026 test set, the pipeline scores 0.3852 on the public leaderboard and 0.4015 on the private leaderboard, versus 0.2714 / 0.2734 for Baseline A (mid-clip time, frame centre) and 0.3107 / 0.3188 for Baseline B (quarter-clip time, frame centre).
  • Per-component scores: The full pipeline achieves classification C = 0.5057, temporal T = 0.3689, and spatial S = 0.3498. Both baselines share C = 0.5807 and S = 0.2505, with T = 0.1896 (Baseline A) and T = 0.2664 (Baseline B). The pipeline therefore trades lower classification accuracy for much larger temporal and spatial gains.
  • Temporal windowing matters: Replacing a uniform midpoint time with a PE δ-expanded window improves the private score from 0.3592 to 0.4015, an improvement of 0.039 over the PE Top-1 frame variant (0.3627 public/private pair: 0.3435/0.3627), suggesting the single highest-similarity frame is a noisy temporal estimate.
  • Prompt ensembling gives smaller but consistent gains: Moving from a 1-prompt structured setting (0.3801 / 0.3961) to 3-prompt majority vote (0.3809 / 0.3978) to 5-prompt plus tiebreaker (0.3849 / 0.4001) to entropy-gated pairwise adjudication (0.3852 / 0.4015), a total improvement of 0.0054 over the single-prompt setting.
  • Spatial grounding is the largest single driver: OWL-v2 with type and scene conditioning raises the score by 0.053 over the centre-of-frame baseline (0.3358 / 0.3487), whereas the Molmo2 pointing alternative scores only 0.2589 / 0.2647.
  • Stage contributions are uneven: Spatial localization contributes most, temporal localization second, and the classification ensemble the least.
  • Cascade dispatch statistics: Out of n = 2027 test clips, 988 (48.7%) are resolved by the base ensemble with no escalation, 305 (15.0%) require the tiebreaker prompt p5, and 734 (36.2%) require full pairwise adjudication.
  • Dominant failure modes: Distant collisions where OWL-v2 returns low-confidence detections on a foreground vehicle, adverse weather or lighting that degrades PE similarity and shifts the temporal window, and shallow-angle rear-end and sideswipe crashes that stay ambiguous even after pairwise adjudication.

Methodology in Plain English

The system answers three questions in sequence rather than all at once.

When. Frames are sampled at 8 FPS and scored by how semantically similar they are to the text query "traffic accident" using Meta's Perception Encoder, a CLIP-style contrastive model. The top-5 similarity peaks define candidate frames; the temporal window spans from the earliest peak minus 2 seconds to the latest peak plus 2 seconds. The predicted accident time is simply the midpoint of that expanded window, so only the earliest and latest selected peaks influence the output.

What. The selected key frames plus metadata (scene layout, weather, time of day, video quality) are passed to Qwen-3.5-VL 9B with five different structured prompts, each stressing a different kind of evidence: direct classification, vehicle motion at impact, contact geometry and angle, contrastive elimination of wrong categories, and a tiebreaker. The five votes are summarized by top-two margin and normalized entropy. If the margin exceeds 2 or entropy falls at or below 0.75, the plurality class wins outright. Otherwise the system escalates: first the tiebreaker prompt adds a vote over at most six evenly spaced key frames, and if ambiguity persists a pairwise adjudicator is restricted to the top-two classes and conditioned on impact geometry and contact point, overriding the plurality if it picks one of those two.

Where. An open-vocabulary detector, OWL-v2, is queried not with a generic phrase but with a prompt built from the predicted accident type (for example, "car crashing into back of another car" for rear-end) suffixed with a scene phrase ("on a highway", "at a signalized intersection", and other layout options). Detections above a 0.05 confidence threshold from all key frames are pooled, the top 5 are kept, and the impact point is the score-weighted centroid of their box centres, with the impact region being the union of those boxes clamped to the frame.

All processing runs on a single NVIDIA L4 GPU with 24 GB of memory using only open-weight models, no fine-tuning, and no proprietary APIs.

Why This Matters

The work shows that decomposing a hard multimodal task into temporal, semantic, and spatial sub-problems can beat end-to-end prompting, and it does so without any labelled real-world training data — a meaningful result for domains where accidents are rare and expensive to annotate. It also quantifies a foreground bias in existing pointing models: OWL-v2 conditioned on accident type and scene outperforms Molmo2 pointing and centre-of-frame heuristics by a wide margin on impact localization.

Real-world applications:

  • Emergency response: faster triage of CCTV feeds, with the predicted impact location helping dispatchers direct responders.
  • Insurance assessment: automated first-pass reconstruction of when, what, and where for claims review.
  • Fleet safety: post-incident analysis and driver coaching on dashcam and depot footage without custom model training per site.
  • Autonomous driving and road-safety analytics: understanding collision geometry from fixed infrastructure cameras to inform risk models and intersection design.

Industry relevance: The method runs on a single 24 GB GPU with 4-bit quantized open-weight models, meaning it is deployable at the edge or on modest server hardware without licensing third-party APIs — a practical constraint for surveillance and fleet operators processing continuous video.

Future Directions

  • Improving distant-collision localization: the dominant failure mode is OWL-v2 selecting a foreground vehicle or missing a far-away impact, so better type-conditioned grounding or multi-scale detection is an obvious next target.
  • Robustness to adverse capture conditions: rain, night, and occlusion degrade the PE similarity scores that drive temporal windowing, raising the question of whether the temporal stage should incorporate motion or weather-aware cues.
  • Resolving shallow-angle ambiguity: rear-end and sideswipe crashes from shallow camera angles remain hard even after pairwise adjudication, suggesting richer geometric reasoning or additional camera viewpoints may be needed.
  • Improving classification without sacrificing T and S: the pipeline's classification score (0.5057) is below both rule-based baselines (0.5807), so the prompt ensemble and adjudication design have clear headroom.

Target Audience

Researchers and practitioners working on video-language models, surveillance analytics, and traffic-safety computer vision; engineers building zero-shot perception pipelines who want a concrete example of task decomposition, prompt ensembling with uncertainty gating, and open-vocabulary detection with multi-frame aggregation; and challenge participants in the ACCIDENT@CVPR benchmark looking for a reproducible baseline that uses only open-weight models on modest hardware.

Authors’ abstract

In this paper, we address the problem of zero-shot understanding of accidents from surveillance videos by identifying when an impact event occurs, what type of impact it is, and where in the frame it occurs using natural language. We propose a three-stage pipeline that decomposes the accident understanding into when, what, and where. The first stage extracts a short temporal window around the impact using vision-language similarity. In the second stage, we perform metadata-driven multi-prompt reasoning with five complementary views (baseline, motion, geometry, contrast, and tiebreaker) and resolve disagreement via an entropy-gated pairwise adjudicator. Finally, we localize the impact of an open-vocabulary detector queried on the predicted accident type and scene layout, and aggregate detections across keyframes using a score-weighted centroid. Our pipeline achieves a substantial improvement in the harmonic-mean score over a centre-of-frame baseline on the zero-shot ACCIDENT @ CVPR benchmark. We show that decomposing zero-shot video understanding into temporal localization, semantic classification, and spatial grounding enable more reliable reasoning with vision-language models than direct prompting alone.

Read the original paper