Skip to content
AI.info

Research

Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding

Overview Research area: Spatio-temporal video grounding (STVG), a computer vision and natural-language task that localizes a described object or event in both space (bounding boxes) and time (a tempor

Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding
arXiv
2610.06018
Published
2026-10-05
Authors
Eryk Kołodziejczyk, Alberto Presta, Karol Szurkowski, Michal Byra

AI summary

Overview

Research area: Spatio-temporal video grounding (STVG), a computer vision and natural-language task that localizes a described object or event in both space (bounding boxes) and time (a temporal segment) within a video.

Technical level: Intermediate. The core idea is intuitive, but the evaluation metrics (m_vIoU, m_tIoU) and the dataset analysis assume some familiarity with video grounding benchmarks.

Scope: The paper audits two leading STVG models on two standard datasets to show that they produce plausible spatio-temporal predictions even when given queries that are unrelated to the video, or no text at all.

What This Paper Is About

STVG models are trained and tested on the assumption that every text query actually describes something in the input video, so the architecture is forced to always output a temporal segment and a bounding box. This paper asks what happens when that assumption is broken: it feeds state-of-the-art STVG models queries taken from unrelated videos, queries describing scenarios absent from the dataset entirely, and no text input whatsoever. The goal is to expose query-insensitive behavior, characterize the false positives it produces, and trace how regularities in the HCSTVG-v2 and VidSTG datasets may encourage it.

Key Contributions

  1. Demonstrates that fully supervised STVG models (TubeDETR and TA-STVG) continue to output plausible spatio-temporal tubes under irrelevant in-domain and out-of-domain queries, with scores computed against the ground-truth annotation of the original positive query.
  2. Provides an in-depth characterization of the HCSTVG-v2 and VidSTG datasets, covering temporal annotation patterns, spatial bounding-box distributions, and query word and verb frequencies, to identify biases that could drive query-insensitive behavior.
  3. Trains TA-STVG variants with empty queries and with the text-processing modules removed entirely, showing how much grounding performance survives without textual guidance.
  4. Introduces a query-blind center-prior baseline to quantify how much residual performance can be explained by dataset-level annotation statistics alone.

Main Findings

  • Models ground irrelevant queries. In Table 1, on HCSTVG-v2, TA-STVG with positive queries scores 0.335 m_vIoU and 0.548 m_tIoU, while negative in-domain queries score 0.206 and 0.476 (a drop of only 39% and 13%), and negative out-of-domain queries score 0.186 and 0.457 (44% and 17%). TubeDETR on the same dataset goes from 0.343 / 0.525 on positive queries to 0.170 / 0.416 on in-domain negatives (50% / 21%) and 0.161 / 0.398 on out-of-domain negatives (53% / 24%).
  • The effect is dataset-dependent. Drops are far larger on VidSTG. For declarative queries, TA-STVG falls from 0.342 m_vIoU and 0.517 m_tIoU to 0.097 and 0.314 (72% / 39%) on in-domain negatives; for interrogative queries it falls from 0.291 / 0.501 to 0.101 / 0.323 (65% / 36%). The paper states the query-insensitive effect is more pronounced on HCSTVG-v2.
  • Temporal grounding is more affected than spatial grounding. The relative gap between positive and negative queries is consistently smaller for m_tIoU than for m_vIoU, which the authors read as a larger problem in temporal localization than in spatial grounding.
  • Predicted boxes tend to hit the main entities. The authors observed that for negative cases the outputted bounding boxes usually corresponded to main entities in the video.
  • Training without real text costs less than expected. In Table 2, TA-STVG trained on empty queries on HCSTVG-v2 drops from 0.357 to 0.262 m_vIoU (27%) and from 0.535 to 0.485 m_tIoU (9%). On VidSTG the same comparison drops from 0.267 to 0.151 m_vIoU (43%) and from 0.468 to 0.399 m_tIoU (15%).
  • Removing text modules entirely still leaves substantial grounding. In Table 3, TA-STVG on HCSTVG-v2 with text-processing modules removed scores 0.238 m_vIoU and 0.479 m_tIoU, versus 0.357 and 0.535 for the baseline.
  • Dataset priors explain only part of the effect. The query-blind center-prior baseline in Table 4 reaches 0.063 m_vIoU and 0.248 m_tIoU on HCSTVG-v2, and 0.014 m_vIoU and 0.089 m_tIoU on VidSTG. Because the empty-query model still processes the video and scores far higher, the authors conclude the residual performance reflects video-conditioned visual cues as well as annotation statistics.
  • Dataset biases are strong. VidSTG has nearly ten times more samples than HCSTVG-v2 and clips ranging from 3 to 180 seconds, while all HCSTVG-v2 samples last 20 seconds. HCSTVG-v2 annotated moments usually span 20–40% of the video; VidSTG moments most often cover either a short 2.5–10% fragment or the entire video. VidSTG moment centers concentrate around the video center. Both datasets show a pronounced vertical bias with bounding-box centers near the frame center; VidSTG is also centrally biased horizontally, while HCSTVG-v2 is more uniform horizontally.
  • Query vocabulary differs between datasets. Using NLTK v. 3.9.1, HCSTVG-v2 is dominated by "the" and "man" with "woman" appearing significantly less often; VidSTG has more "adult" and introduces "children" and "babies". HCSTVG-v2 verbs are action-oriented ("turns", "walks", "takes", "puts", "looks"), whereas VidSTG frequently uses "is", reflecting more descriptive queries.

Methodology in Plain English

The researchers took two fully supervised STVG models, TubeDETR and the newer TA-STVG, and used the authors' released checkpoints for the best-performing configurations. They built two kinds of negative queries. In-domain negatives were existing queries from the same dataset reassigned to a different, unrelated video, which keeps the wording plausible but breaks the query-video link. Out-of-domain negatives came from the query set released by Flanagan et al. (2025), generated from scenarios unlikely to appear in these datasets, such as physics laboratories and mathematics classes. Negative queries were used only at inference time, and scores were always computed against the ground-truth annotation of the original positive query, so a high score under a negative query is a false positive.

To separate the role of text from the role of visual and dataset cues, they then retrained TA-STVG with empty queries, retrained it with the text-processing components removed on HCSTVG-v2, and compared both to a query-blind baseline that always predicts a fixed centered box and a fixed centered temporal segment derived from training-set statistics. They also profiled the two datasets directly, plotting moment duration, moment midpoints, average bounding-box centers and heatmaps, and running word-frequency analysis on the text.

Why This Matters

Impact on research. The paper argues that standard STVG evaluation, which only measures localization accuracy on positive query-video pairs, hides a systematic failure mode. It motivates negative-aware evaluation protocols and architectures that explicitly assess whether a query is relevant before grounding it.

Real-world applications:

  • Video search and editing tools that take natural-language queries, where a user query with no match should return nothing rather than a confidently wrong clip.
  • Surveillance and monitoring systems that localize described events, where false positives on irrelevant descriptions waste operator attention.
  • Video assistants and accessibility tools that answer "find the moment where X happens", where silently grounding an absent X is worse than admitting no match.
  • Robotics or embodied agents that follow language instructions grounded in video, where acting on an irrelevant instruction is a safety concern.

Industry relevance. The work comes from Samsung AI Center, Warsaw, and the failure mode it documents directly affects deployed video understanding products. The finding that TA-STVG retains 0.479 m_tIoU on HCSTVG-v2 with its text modules removed indicates that reported benchmark gains may partly reflect dataset regularities rather than genuine language grounding, which matters for anyone selecting or benchmarking models.

Future Directions

  • Develop more diverse STVG datasets with less centralized temporal and spatial annotation structure, so models cannot lean on dataset priors.
  • Add query-validity estimation to STVG systems, so a model can confirm the absence of a described event instead of being forced to output a tube.
  • Explore alternative negative-query generation strategies: random reassignment, semantic filtering, LLM-generated out-of-domain negatives, and manually curated negatives, each with its own trade-offs in scalability and reliability.
  • Design architectures or training objectives, such as contrastive learning, that increase a model's dependence on textual input rather than on salient visual content alone.

Target Audience

Researchers and engineers working on video-language understanding, video moment retrieval and temporal grounding, and multimodal model evaluation. It is also relevant to practitioners who deploy video search or event-localization systems and need to know how these models behave on queries with no valid answer, and to benchmark designers interested in negative-aware evaluation protocols.

Authors’ abstract

Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.

Read the original paper