Skip to content
AI.info

Research

VideoSTF: Stress-Testing Output Repetition in Video Large Language Models

Overview Research area: multimodal vision-language evaluation, specifically generation stability in Video Large Language Models (VideoLLMs). Technical level: Advanced. Scope: The paper introduces Vide

VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
arXiv
2602.10639
Published
2026-02-11
Authors
Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong

AI summary

Overview

Research area: multimodal vision-language evaluation, specifically generation stability in Video Large Language Models (VideoLLMs). Technical level: Advanced. Scope: The paper introduces VideoSTF, a framework of three n-gram-based metrics, a 10,000-video testbed, and a library of temporal transformations, to measure, stress-test, and adversarially induce output repetition across 10 VideoLLMs.

What This Paper Is About

VideoLLMs are usually judged on whether their answers are correct or factual, but the authors show they can also fail in a different way: the model degenerates into a self-reinforcing loop that repeats the same phrases or sentences until it hits its output token limit. The paper's goal is to define this failure precisely, build a standard way to measure it, and test how easily it can be triggered by disturbing the temporal ordering of a video's frames.

Key Contributions

  1. The paper identifies output repetition as a distinct generation-stability failure in modern VideoLLMs, separate from accuracy, factual error, or hallucination benchmarks, and argues existing evaluation protocols miss it.
  2. It proposes VideoSTF, described as the first framework specifically for measuring and stress-testing output repetition in video-language generation, with three n-gram-based metrics (Repetition Rate, Repetition Intensity, Information Entropy).
  3. It provides a standardized testbed of 10,000 videos sampled from public video instruction datasets (LLaVA-Video-178K, NeXT-QA, ActivityNetQA, LLaVA-Hound) with durations up to 180 seconds, plus a temporal stressor library of five controlled transformations (Add, Delete, Replace, Reverse, Shuffle).
  4. It demonstrates through experiments on 10 representative VideoLLMs that repetition is pervasive, strongly amplified by temporal transformations even when semantic content is preserved, and efficiently inducible as a black-box attack.

Main Findings

  • Repetition is widespread across models and base LLMs. All 10 tested VideoLLMs exhibit output repetition despite being built on different underlying LLMs. ShareGPT4Video and Molmo2-8B are the most severe, with repetition rates exceeding 79% and 65% respectively. Specifically, ShareGPT4Video records RR of 91, 85, 79, and 82 at 8, 16, 24, and 32 sampled frames, and Molmo2-8B records 65, 69, 66, and 72.
  • Frame count does not drive the failure. Repetition remains stable as the number of sampled frames varies across 8, 16, 24, and 32, indicating the failure mode is insensitive to input temporal length. A small subset of models, such as LLaVA-Video-7B-Qwen2 (RR of 3, 5, 10, 7 across the four frame settings), shows increasing repetition with more frames, suggesting denser sequences of visually similar frames make them more susceptible.
  • Recurring visual content triggers looping language. Videos containing recurring or highly similar scenes are more prone to repetition, and outputs often lock onto characteristic looping phrases such as "continues to".
  • Temporal transformations amplify repetition. Repetition rates under transformation are higher than on original videos in almost all cases. Mean repetition rate increases by 78% to 205% for LLaVA-Video-7B-Qwen2 and by 67% to 108% for Qwen3-VL-8B-Instruct, with in extreme cases repetition rates over 90%. The same trend appears in Repetition Intensity and Information Entropy.
  • Transformation structure matters. Add, Delete, and Replace preserve partial temporal coherence while introducing redundancy or localized inconsistency, and are particularly harmful. Reverse globally disrupts temporal order and reduces the tendency to lock into repetition. Shuffle sits between the two extremes.
  • Repetition is an exploitable black-box attack. Under a threat model where the attacker can modify the sampled video frames but cannot access model internals, attacks induce repetition with only tens of queries. Highest Attack Success Rate reaches 98%, and the maximum Average Queries is 15.8. Reverse is the weakest attack (AQ of 1.0 across reported settings), while models with low original repetition remain vulnerable: VideoLLaMA2 achieves 40% to 83% ASR and InternVL3.5-8B achieves 28% to 72% ASR.
  • Some models break down at high frame counts. LLaVA-NeXT-Video-7B and LLaVA-NeXT-Video-7B-DPO produce empty outputs at 32 sampled frames because visual embeddings exceed the underlying LLMs' maximum input capacity. When the sampled frame number falls in [28, 32), some videos still yield empty outputs while others produce highly repetitive responses that often contain grammatical errors.
  • Metric choice affects sensitivity. Testing n from 1 to 10 shows smaller n (n ≤ 2) yields higher RR but overcounts common short phrases, while larger n increasingly fails to capture repetition, especially at n ≥ 7. For RI and IE, repetition becomes difficult to observe when n ≥ 4. Relative cross-model repetition tendencies remain stable across n, with ShareGPT4Video, LLaVA-Video-7B-Qwen2-Video-Only, and Molmo2-8B consistently higher.

Methodology in Plain English

The authors treat a video as an ordered sequence of sampled frames fed to a VideoLLM, which then generates text token by token. To detect repetition, they split each output into overlapping n-grams (contiguous token sequences of length n) and apply three measures. Repetition Rate (RR) counts the fraction of outputs where any n-gram appears more than once, using n = 5 to avoid flagging ordinary function words. Repetition Intensity (RI) follows the Rep-n formulation and measures the proportion of n-grams that are duplicates, using n = 1. Information Entropy (IE) measures the normalized n-gram entropy of the output, where lower entropy means less lexical diversity, also at n = 1.

For testing, they assembled 10,000 videos from several public video instruction datasets, spanning short clips to videos up to 180 seconds and covering everyday categories such as comedy, lifestyle, and sports. They then defined five temporal transformations that change the ordering or composition of frames while largely preserving semantic content: Add (insert k random frames), Delete (remove k random frames), Replace (swap k frames with other frames from the same video), Reverse (invert frame order), and Shuffle (randomly permute frames). For Add, Delete with k = 2, and Replace, stochastic choices were run for 30 trials; Shuffle used 30 random permutations per video; deleting one frame was done exhaustively over all sampled frames.

Three experiments were run. Pervasive testing measured repetition on natural videos at 8, 16, 24, and 32 frames. Temporal stress testing applied the transformations and compared repetition against the originals. Adversarial exploitation selected videos that originally produced non-repetitive outputs and iteratively applied transformations, querying the target model until repetition appeared or a maximum of 30 query attempts was reached, reporting Attack Success Rate (ASR) and Average Queries (AQ). All 10 models were run deterministically with do_sample set to False and temperature fixed at 0.0.

Why This Matters

The paper reframes repetition from a cosmetic quality issue into a reliability and security concern. It argues that stability of autoregressive generation is an implicit assumption in VideoLLM research that current benchmarks do not test, and that evaluation should become stability-aware. Because repetition can be induced with few black-box queries and may run until the output token limit, it creates a practical denial-of-service surface.

Real-world applications affected:

  • Deployed video captioning and video summarization services, where looping outputs waste compute and produce unusable descriptions.
  • Video question answering and video assistants, where repetitive degeneration makes responses unhelpful to end users.
  • Accessible video description and content indexing pipelines that depend on reliable, diverse generated text.
  • Any hosted video understanding API, where an attacker with only input access could induce runaway generation and exhaust serving capacity.

Industry relevance: organizations serving VideoLLM inference at scale bear both the compute cost of degenerate generation and the security risk of deliberate triggering, so the metrics and testbed offer a way to audit and compare model robustness before deployment.

Future Directions

  • Developing mitigation and stabilization mechanisms that explicitly model or regularize temporal redundancy, going beyond decoding heuristics. The paper states that mitigation is beyond its scope.
  • Characterizing why models lock into self-reinforcing loops, since the results link repetition to aggregation and attention over sequences of visually similar frames.
  • Extending stability-aware evaluation to broader settings, including combinations of temporal transformations and other modalities, and incorporating repetition metrics into standard VideoLLM benchmarks alongside accuracy and hallucination measures.
  • Studying defenses against the black-box, input-level attack surface, and whether the reported pattern of Add, Delete, and Replace being more damaging than Reverse or Shuffle generalizes to other model families and larger testbeds.

Target Audience

Researchers and engineers working on multimodal large language models, video-language systems, and model evaluation; benchmark designers who need metrics beyond task accuracy; and AI safety and security practitioners interested in the reliability and attack surface of deployed video understanding systems.

Authors’ abstract

Video Large Language Models (VideoLLMs) have recently achieved strong performance in video understanding tasks. However, we identify a previously underexplored generation failure: severe output repetition, where models degenerate into self-reinforcing loops of repeated phrases or sentences. This failure mode is not captured by existing VideoLLM benchmarks, which focus primarily on task accuracy and factual correctness. We introduce VideoSTF, the first framework for systematically measuring and stress-testing output repetition in VideoLLMs. VideoSTF formalizes repetition using three complementary n-gram-based metrics and provides a standardized testbed of 10,000 diverse videos together with a library of controlled temporal transformations. Using VideoSTF, we conduct pervasive testing, temporal stress testing, and adversarial exploitation across 10 advanced VideoLLMs. We find that output repetition is widespread and, critically, highly sensitive to temporal perturbations of video inputs. Moreover, we show that simple temporal transformations can efficiently induce repetitive degeneration in a black-box setting, exposing output repetition as an exploitable security vulnerability. Our results reveal output repetition as a fundamental stability issue in modern VideoLLMs and motivate stability-aware evaluation for video-language systems. Our evaluation code and scripts are available at: https://github.com/yuxincao22/VideoSTF_benchmark.

Read the original paper