Skip to content
AI.info

Research

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

Overview Research area: Computer Vision / multimodal video-language understanding; specifically, evaluation benchmarks for narrative comprehension in video. Technical level: Intermediate. The ideas ar

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
arXiv
2601.01095
Published
2026-01-03
Authors
Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma, Gargi Chakraborty

AI summary

Overview

  • Research area: Computer Vision / multimodal video-language understanding; specifically, evaluation benchmarks for narrative comprehension in video.
  • Technical level: Intermediate. The ideas are conceptually accessible, but the paper assumes familiarity with multimodal large language models (MLLMs), object detection, re-identification tracking, and benchmark design.
  • Scope in one sentence: The paper introduces NarrativeTrack, a 1,006-question benchmark built from 406 video clips (average 55.3 seconds) that tests whether MLLMs can track individual human entities through a three-stage "Compositional Reasoning Progression" — entity existence, entity changes, and entity ambiguity.

What This Paper Is About

Existing video benchmarks for MLLMs mostly test either short-clip recognition or coarse whole-video summarization, so they often can be answered from a single frame or from language priors rather than genuine reasoning over time. The authors argue that narrative understanding actually requires tracking who is present, what they are doing, and how their state changes across scene transitions and temporal gaps. NarrativeTrack is their answer: a benchmark, plus a fully automated pipeline for building it, that evaluates models through fine-grained, entity-centric questions about the people in a video.

Key Contributions

  1. NarrativeTrack, a new entity-centric narrative benchmark. It contains 1,006 QA pairs drawn from 406 video clips, with an average clip length of 55.3 seconds and temporal scales up to 659 seconds, covering genres such as documentary, news, and TV drama. The paper describes it as the first benchmark to evaluate narrative understanding in MLLMs through fine-grained entity-centric reasoning.

  2. The Compositional Reasoning Progression (CRP). A structured evaluation framework that increases complexity across three dimensions: entity existence (temporal continuity), entity changes (actions, outfits, scenes), and entity ambiguity (fine-grained perceptual disambiguation among visually similar entities). The final question distribution is entity existence (200), action changes (222), outfit changes (192), scene changes (191), and entity ambiguity (201).

  3. A fully automated entity-centric pipeline. Three stages — entity detection, entity tracking, and contextual recognition — extract temporally grounded entity representations (bounding box trajectory, action, scene, outfit) directly from raw video without human supervision.

  4. A broad empirical study of 13 open-source and 7 proprietary MLLMs, with analysis of temporal directional bias, entity continuity types, distractor types, and frame density.

Main Findings

  • Proprietary models lead by a wide margin. Gemini-2.5-Pro achieves the highest average accuracy of 83.80%, followed by Gemini-3.1-Pro at 83.40%. Open-source models remain far behind.

  • General-purpose models beat video-specialized ones among open-source MLLMs. Qwen-2.5-VL-32B reaches 56.96%, compared with 49.70% for the best video-specialized model, Video-LLaMA2-72B.

  • A trade-off between perceptual grounding and temporal reasoning. General-purpose MLLMs show strong perceptual grounding but weak temporal continuity; video-specialized MLLMs capture temporal context but frequently hallucinate entity contexts.

  • Scaling helps inconsistently. InternVL3-38B outperforms its 8B version by +6.66%, and Video-LLaMA2-72B outperforms its 7B version by +5.66%, yet LLaVA-NeXT-Video-34B slightly underperforms its 7B version.

  • Frame density does not guarantee improvement. Scaling temporal coverage did not produce performance gains in the authors' analysis.

  • Results depend on genuine multimodal and temporal grounding. Removing visual inputs reduces GPT-4o accuracy by 30.52%, approaching the random baseline. Reversing video frames drops performance on the entity change ordering task from 51.2% to 6.1%.

  • Temporal directional bias. Models reason better forward than backward. The forward-versus-backward gap is 20.65% for open-source general-purpose MLLMs, 9.96% for open-source video-specialized MLLMs, and 17.55% for proprietary MLLMs. The bidirectional "agnostic" ordering condition yields the lowest accuracy.

  • Different strengths on entity continuity. General-purpose MLLMs perform best on disappear cases, relying on static visual cues; video-specialized and proprietary MLLMs do better on reappear cases, showing stronger temporal integration but a higher tendency to hallucinate.

  • Synthetic distractors expose hallucination. General-purpose MLLMs show stable performance across real and synthetic distractors, whereas video-specialized and proprietary MLLMs drop notably on synthetic distractors.

  • Pipeline quality. Ensemble detection improves recall to 0.848 from 0.780 (Detectron2 alone) and 0.801 (OWLv2 alone). Model-based tracking verification agrees with human majority vote on 96.08% of cases across 1,108 detections.

  • Dataset quality. In a triple-annotator review of 100 randomly sampled QA pairs, 70% were unanimously judged valid (Fleiss' κ = 0.767). After cleanup, three annotators re-evaluated the benchmark and achieved 96% average human accuracy.

Methodology in Plain English

The authors start from raw video and build structured "entity cards" for the people in it. Detection runs two off-the-shelf detectors (Detectron2 and OWLv2), and boxes from the two are merged only when they overlap enough (IoU ≥ 0.5), keeping the higher-confidence box. Those per-frame boxes are then linked across time into trajectories; an OSNet-x1.0 re-identification model clusters them into identities, and the four largest clusters are treated as the main characters. Trajectories are refined with face recognition and ensemble verification across multiple MLLMs using majority voting. Finally, Gemini-2.5-Pro labels each timestep with the entity's action, outfit, and scene, using clips where the target entity is highlighted.

Questions are then generated programmatically by filling fixed templates with these extracted attributes, with ground-truth answers taken from the same structured data. GPT-4o is used only for grammar and fluency cleanup and to filter invalid cases. Three answer formats are used — binary, multiple-choice, and ordering — and three reasoning patterns (forward, backward, agnostic) are produced by varying the temporal reference point. Distractors are either real (other entities in the same clip) or synthetic (entities from different clips). Scoring is deterministic exact-match accuracy, with no LLM judging.

Why This Matters

Impact on research. The paper reframes narrative understanding as a compositional capability requiring both perceptual precision and temporal reasoning, and provides a diagnostic that separates failure modes: broken temporal continuity, weak grounding of evolving attributes, and confusion between similar entities. The finding that models can "extend a narrative but cannot rewind it" — with the reversal of frame order cutting ordering accuracy from 51.2% to 6.1% — gives the community a concrete, measurable target for architectural work.

Real-world applications:

  • Long-video assistants that must answer "what happened to this person" across a full recording rather than a short clip.
  • Accessibility tools that generate audio descriptions or recaps of films and TV episodes, where characters change clothes, locations, and actions.
  • Media and archive search, letting editors or journalists locate a specific participant's appearances and state changes across footage.
  • Security and compliance review of long recordings, where the question is whether a specific individual persisted, disappeared, or reappeared.

Industry relevance. The benchmark is released with code and data, and its pipeline is fully automated and built on off-the-shelf components, so organizations with large video corpora can generate entity-level evaluation data and training signal without manual labeling. The measured trade-off between perceptual grounding and temporal reasoning directly informs how video-capable models are trained and which base model families are better suited to narrative tasks.

Future Directions

  • Bidirectional and reversal-aware temporal modeling. The authors leave architectural innovations such as bidirectional temporal modeling and contrastive reversal objectives — which would enforce symmetry between forward and backward reasoning over entity states — as future work.
  • Extending beyond human entities. The current benchmark is restricted to visually salient human subjects; extending the framework to non-human entities is described as a promising direction.
  • Closing the perceptual-temporal gap. The results suggest narrative understanding emerges only from integrating temporal reasoning with perceptual precision, but the paper does not propose a model that achieves this integration.
  • Understanding why frame density fails to help. The observation that scaling temporal coverage does not guarantee performance gains is reported but not resolved.

Target Audience

Researchers and engineers working on multimodal large language models, video understanding, and benchmark design; practitioners building long-video assistants, media analysis tools, or accessibility systems; and anyone interested in how model evaluation can be made more diagnostic by decomposing a capability into progressively harder reasoning levels.

Authors’ abstract

Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative understanding requires grounding who is doing what, when, and where, maintaining coherent entity representations across dynamic visual and temporal contexts. We introduce NarrativeTrack, the first benchmark to evaluate narrative understanding in MLLMs through fine-grained entity-centric reasoning. Unlike existing benchmarks limited to short clips or coarse scene-level semantics, we decompose videos into constituent entities and examine their continuity via a Compositional Reasoning Progression (CRP), a structured evaluation framework that progressively increases narrative complexity across three dimensions: entity existence, entity changes, and entity ambiguity. CRP challenges models to advance from temporal persistence to contextual evolution and fine-grained perceptual reasoning. A fully automated entity-centric pipeline enables scalable extraction of temporally grounded entity representations, providing the foundation for CRP. Evaluations of state-of-the-art MLLMs reveal that models fail to robustly track entities across visual transitions and temporal dynamics, often hallucinating identity under context shifts. Open-source general-purpose MLLMs exhibit strong perceptual grounding but weak temporal coherence, while video-specific MLLMs capture temporal context yet hallucinate entities' contexts. These findings uncover a fundamental trade-off between perceptual grounding and temporal reasoning, indicating that narrative understanding emerges only from their integration. NarrativeTrack provides the first systematic framework to diagnose and advance temporally grounded narrative comprehension in MLLMs.

Read the original paper