Skip to content
AI.info

Research

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models Overview Research area: Computer Vision, specifically content-based video retrieval (CBVR), video surv

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
arXiv
2601.04824
Published
2026-01-08
Authors
Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera

AI summary

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

Overview

  • Research area: Computer Vision, specifically content-based video retrieval (CBVR), video surveillance, and Multimodal Large Language Models (MLLMs).
  • Technical level: Intermediate. The paper assumes familiarity with retrieval metrics (mAP), contrastive vision-language models, sentence embeddings, and multimodal prompting, but the core ideas are presented without heavy mathematical machinery.
  • Scope: The paper introduces SOVABench, a retrieval benchmark built from real vehicle-surveillance footage and organized around opposite action pairs, and proposes a training-free framework that turns MLLM-generated text into embeddings for retrieval and classification.

What This Paper Is About

Existing content-based video retrieval benchmarks mostly measure scene-level similarity and do not test whether a system can tell apart actions that look alike, such as loading versus unloading a vehicle. The authors build SOVABench from real surveillance footage to evaluate exactly this action discrimination and temporal-direction understanding, using two protocols (inter-pair and intra-pair). Alongside the benchmark, they present a training-free "MLLM-to-Embedding" pipeline that asks a Multimodal Large Language Model to describe an image or video, then converts the resulting sentences into embeddings for similarity comparison.

Key Contributions

  1. SOVABench benchmark. The authors state it is the first content-based video retrieval benchmark built from real-world vehicle-surveillance footage, designed to evaluate action discrimination and temporal direction understanding. It is constructed from two surveillance sources, MEVA and the VIRAT validation set, with seven opposite vehicle-action pairs and spatially cropped, temporally aligned clips.
  2. Two evaluation protocols. An inter-pair protocol measures whether models can distinguish between different action pairs, and an intra-pair protocol measures whether models can distinguish temporally inverse actions within a pair (e.g., open versus close).
  3. A training-free, instruction-following embedding framework. The MLLM-to-Embedding framework uses off-the-shelf MLLMs as black-box describers, splits their text into sentences, embeds each sentence, and compares samples using a maximum pairwise cosine similarity.
  4. Validation on image-based tasks plus SOVABench. The framework is first tested on spatial and counting benchmarks where contrastive vision-language models often fail, and then applied to the SOVABench retrieval settings.

Main Findings

  • Inter-pair retrieval is hard but non-trivial for all models. The random baseline on SOVABench (Inter-pair) is 3.4 mAP. All evaluated models considerably exceed it. The highest score is MiniCPM-V 4.5 with the MLLM-to-Embedding framework and task-aware prompting at 38.3 mAP.
  • Contrastive video-VLMs are competitive. CLIP4Clip reaches 36.6 mAP on SOVABench (Inter-pair), the second-best performance reported, followed by VideoCLIP at 34.5 and ActionCLIP at 32.8. Contrastive image-VLMs score lower: CLIP-ViT-L-14 29.1, SigLIP2-Giant 30.6, MERU 28.6.
  • Intra-pair retrieval is close to chance for every model. The random baseline is 50.3 Pair-mAP. The best reported result is MiniCPM-V 4.5 TASK-AWARE at 53.6 Pair-mAP; Gemini 2.5 Flash TASK-AWARE follows at 53.9... which is reported as 53.9 Pair-mAP. Other contrastive baselines sit at 51.1 to 51.4 Pair-mAP. The authors state that all models perform only slightly above random, showing systems struggle to distinguish temporally inverse actions.
  • Task-aware prompting generally helps. In SOVABench, task-aware configurations consistently outperform their general-instruction counterparts. For Inter-pair, MiniCPM-V 4.5 improves from 34.4 to 38.3 mAP and InternVL3.5 8B from 27.7 to 35.4 mAP. Video-MLLMs also gain (VideoChat-R1 7B from 25.1 to 31.6 mAP), as does Gemini 2.5 Flash (27.9 to 33.2 mAP).
  • Video-specialized MLLMs do not hold a consistent advantage. The authors report that Video-MLLMs show no consistent benefit over general-purpose MLLMs on SOVABench, suggesting their temporal modeling may not align with short atomic surveillance actions.
  • Large API models do not necessarily win. Qwen3-VL 235B A22B reaches 14.7 mAP GENERAL and 29.1 mAP TASK-AWARE on Inter-pair, below open-source alternatives; Gemini 2.5 Flash reaches 33.2 mAP TASK-AWARE.
  • The framework beats CLIP on image-based spatial and counting tasks. Averaged over the spatial benchmarks, the best configuration gains 11.5% absolute over CLIP, and 34.2% absolute over CLIP on object counting. CLIP-ViT-B/32 scores 36.5 spatial average, 31.5 count average, and 34.0 overall, while InternVL3.5 8B TASK-AWARE reaches 47.9 spatial, 60.6 count, and 54.3 overall, and MiniCPM-V 4.5 TASK-AWARE reaches 38.8 spatial, 65.7 count, and 52.3 overall.
  • Task-aware prompting splits for spatial understanding. It improves InternVL3.5 8B by +4.6% on spatial tasks but reduces MiniCPM-V 4.5 by -6.6%, whereas for object counting it delivers +13.4% for InternVL3.5 8B and +15.3% for MiniCPM-V 4.5. The authors hypothesize that spatial understanding contains many more sub-tasks, making general instructions less useful.
  • Sentence splitting with maximum aggregation helps overall. Using InternVL3.5 2B, the module improves spatial benchmarks (e.g., INTERNAL GENERAL SpatialBench Indoor 37.1 versus 35.0 without it) but shows a moderate reduction in counting, while overall average performance benefits. It also preserves semantics in arbitrarily long MLLM outputs.
  • Embedder choice matters little. On SOVABench (Inter-pair) with MiniCPM-V 4.5 TASK-AWARE, scores range only from 33.5 mAP (CLIP-ViT-L-14 text tower, 123M parameters, 768 dimensions) to 38.3 mAP (GTE-Large-8152, 409M parameters, 1024 dimensions). No clear relationship between parameter count, embedding size, and performance is observed.
  • Denser frame sampling does not improve SOVABench. With MiniCPM-V 4.5 TASK-AWARE, Inter-pair mAP is 38.3 at 1 FPS, 37.3 at 3 FPS, 36.3 at 5 FPS, and 36.2 at 7 FPS. Intra-pair Pair-mAP is 53.6, 54.0, 53.8, and 53.4 respectively.
  • Failures come largely from text generation, not embedding quality. Error analysis on the Open trunk / Close trunk pair with MiniCPM-V 4.5 TASK-AWARE finds 34 generation errors or hallucinations, 16 temporal misunderstandings, 12 under-descriptions, and 9 action assumptions, with only 26 totally correct samples out of 97.
  • Constrained evaluation favors MLLMs more. When distracting human-only samples are removed and Inter-pair is restricted to the 1,423 queries, random is 23.7 mAP, the best contrastive VLM is VideoCLIP at 41.4 mAP, and the best MLLM configuration is MiniCPM-V 4.5 TASK-AWARE at 44.8 mAP.
  • Task-aware prompting is faster as well as better. In the efficiency analysis (measured on NVIDIA GeForce RTX 3090 GPUs), MLLMs are heavier and slower than contrastive VLMs, but task-aware configurations consistently process more instances per second than their general counterparts (e.g., MiniCPM-V 4.5 TASK-AWARE 0.26 versus GENERAL 0.10).

Methodology in Plain English

The authors start from two existing surveillance datasets, MEVA and the VIRAT validation set, and pull out vehicle-related activities. They group these activities into seven pairs of opposite actions, such as loading versus unloading, entering versus exiting, and turning left versus turning right. For each activity they identify the objects involved and define a spatial region of interest enclosing all actors, producing stable video crops that isolate the action from the background. Each clip is trimmed to the duration of its annotated activity so it captures exactly one action.

Two evaluation setups are built on top of this. In the inter-pair setup, each opposite pair is merged into a single class, giving six query classes (the Drive forward / Reverse pair is excluded because those motions always co-occur with other movement actions). Each queried clip is retrieved against all 9,882 samples in the benchmark, which includes human-only surveillance activities as distractors, and the metric is mean Average Precision. In the intra-pair setup, each opposite pair becomes its own binary retrieval problem where only the opposite action counts as non-relevant; the metric is Pair-mAP, the average of the mAP values across pairs.

For the retrieval method, the authors avoid training anything. They give a vision-language model a visual input plus a textual instruction, either a general "Describe the image/video" or a task-aware prompt that names the type of information to extract. The model's text response is split into lines and then into sentences using the NLTK sentence splitter, and each sentence is encoded separately with a sentence-similarity text encoder. The similarity between two samples is the maximum cosine similarity between any pair of their sentence embeddings, on the reasoning that one discriminative sentence about a motion cue can be enough. They use GTE-Large-8152 as the default sentence encoder and greedy decoding for the MLLM generation step. Task-aware prompts never mention class names or evaluation protocols, keeping the setting zero-shot.

Why This Matters

  • Impact on research: The paper isolates a capability that existing CBVR benchmarks do not probe, namely discriminating actions that differ mainly in their temporal direction. By releasing an evaluation protocol, annotations, extraction instructions, and code, it gives the community a way to measure a known weak spot of multimodal models rather than a scene-similarity proxy.
  • Real-world applications:
    • Alarm filtering in video management systems, where operators need to suppress recurring events that are similar to but not the same as an alert condition.
    • Recurrent event and behavior analysis across many cameras, where retrieving "the same kind of thing that happened before" is the core operation.
    • Incident review, where investigators search footage by action rather than by scene appearance.
    • Search interfaces over vehicle-centric footage, such as finding all loading or unloading events in a parking area.
  • Industry relevance: The work is co-authored with Milestone Systems A/S, a video management software company, and the authors explicitly frame the benchmark around alarm filtering and recurrent event detection. The finding that task-aware prompting improves both accuracy and inference speed is directly useful for deployment, since the framework is training-free, works with off-the-shelf models, and can call API models as black boxes. The efficiency table also makes clear that MLLM-based retrieval is much slower than contrastive VLMs, which matters for operational budgets.

Future Directions

  • Improving prompting and embedding mechanisms so that temporal progression and action dynamics are encoded more precisely; the authors name this as the primary next step.
  • Reducing the failure modes identified in the error analysis, particularly hallucinated objects and reversed temporal direction, which the paper shows dominate retrieval errors on SOVABench (Intra-pair).
  • Closing the gap on temporal direction understanding, where every evaluated model sits only slightly above the 50.3 Pair-mAP random baseline.
  • Making MLLM-based retrieval faster, since the efficiency analysis shows MLLMs processing far fewer instances per second than contrastive VLMs, and the authors note that shorter task-aware outputs help but do not eliminate the gap.

Target Audience

Researchers and engineers working on video retrieval, video surveillance analytics, and multimodal large language models. It is particularly relevant to those building retrieval or alarm-filtering systems on surveillance footage, to practitioners comparing contrastive VLMs against MLLM-based pipelines, and to benchmark designers interested in how to construct retrieval evaluations around action semantics and temporal direction rather than scene similarity.

Authors’ abstract

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate the action discrimination required in surveillance. To address this gap, we introduce SOVABench (Surveillance Opposite Vehicle Actions Benchmark), a real-world retrieval benchmark built from surveillance footage and centered on vehicle-related actions. SOVABench defines two evaluation protocols (inter-pair and intra-pair) to assess cross-action discrimination and temporal direction understanding. Although action distinctions are generally intuitive for human observers, our experiments show that they remain challenging for state-of-the-art vision and multimodal models. Leveraging the visual reasoning and instruction-following capabilities of Multimodal Large Language Models (MLLMs), we present a training-free framework for producing interpretable embeddings from MLLM-generated descriptions for both images and videos. The framework achieves strong performance on SOVABench as well as on several spatial and counting benchmarks where contrastive Vision-Language Models often fail. The code, annotations, and instructions to construct the benchmark are publicly available.

Read the original paper