Skip to content
AI.info

Research

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Overview Research area: Computer vision and multimodal long-video understanding, specifically memory architectures for question answering over day-long and week-long video recordings. Technical level:

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
arXiv
2609.38155
Published
2026-09-29
Authors
Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua

AI summary

Overview

Research area: Computer vision and multimodal long-video understanding, specifically memory architectures for question answering over day-long and week-long video recordings.

Technical level: Advanced. The paper assumes familiarity with retrieval-augmented generation, graph-based memory, vision-language models, object tracking, and PageRank-style relevance propagation.

Scope: The paper introduces Grounded Entity Biographies (GEB), a long-video memory framework that links visually grounded observations of the same physical object or person across clips into retrievable "biographies," and evaluates it on four question-answering benchmarks.

What This Paper Is About

Long-video memory systems usually organize a recording as a timeline of events described in text. This makes it hard to tell whether two mentions of, say, "a red mug" refer to the same physical mug, because different objects can share a description and one object can be described differently as its state or location changes. GEB's goal is to resolve physical identity in the memory itself, so that retrieval can follow a single object or person through hours or days of footage rather than only retrieving text that happens to match the question.

Key Contributions

  1. Grounded entity biographies: A memory representation and construction procedure that links visually grounded observations of the same inferred physical instance into a temporally ordered biography, while retaining each observation's event context and supporting frames.
  2. Retrieval and reading through identity: A mechanism that connects episodes through shared physical entities, propagates relevance from a matched observation to the entity's other appearances and their episodic context, and presents biography excerpts that also list not-yet-inspected appearances as targets for further search.
  3. Conservative grounded association: An association rule combining multimodal embedding similarity with a visual separation test on shared frames, where evidence that two subjects appear apart vetoes a match even when their descriptions are similar.
  4. Empirical validation and analysis: Improvements on long-video question answering under matched controller and answer model (Qwen3.5-35B), plus evidence-access diagnostics and ablations isolating grounding, association, and biography reading.

Main Findings

  • EgoLifeQA: GEB reaches 72.0% accuracy, 4.4 percentage points above the best published result, MAGIC-Video at 67.6%. Both use the same retrieval controller, answer model, and retrieval limits.
  • Ego-R1-Bench: GEB averages 71.3% versus MAGIC-Video's 64.7%, a 6.6 point gain. Results are averaged over three seeds.
  • Per-family gains: GEB improves over MAGIC-Video on four of the five EgoLifeQA question families, with the largest gain in EventRecall (+7.9 points). Family scores: EntityLog 70.4, EventRecall 73.8, HabitInsight 73.8, RelationMap 71.2, TaskMaster 71.4.
  • MM-Lifelong open-ended answering: On Test@Week GEB reaches 36.83%, exceeding the strongest competing result, WorldMM at 31.42%, by 5.41 points. On the gameplay Test@Day split it reaches 17.58%, a 0.83 point lead over ReMA.
  • Evidence access: The fraction of EgoLifeQA questions whose evidence window is reached rises from 37.6% to 58.9% over MAGIC-Video, across all question families.
  • Distributed evidence: On MultiHop-EgoQA, complete evidence coverage rises from 35.4% to 52.1% over MAGIC-Video, a 47% relative gain, with larger relative gains on questions requiring three or more intervals.
  • Temporal grounding: GEB's answer score on MultiHop-EgoQA increases from WorldMM's 2.74 to 3.11, while mIoU increases by 1.6 points (18.9 to 20.5). GEB exceeds every model the benchmark reports in IoU@0.3 and mIoU, including GeLM, which is fine-tuned on the benchmark's training set.
  • Not just more retrieval: Raising MAGIC-Video's MultiHop-EgoQA allowance from three to six units per round leaves complete coverage below GEB at similar retrieved duration. WorldMM achieves comparable complete coverage to GEB (0.526 vs 0.521) but its retrieved units span more of the clip (Clip 0.399 vs 0.343) and its answer score is lower.
  • Ablations on EgoLifeQA (full GEB = 72.0%): Removing association drops to 68.6% (−3.4); keying identity by described name drops to 69.2% (−2.8); appending descriptions to captions drops to 68.2% (−3.8); removing observation-to-timeline edges drops to 68.2% (−3.8); removing same-instance edges costs 3.0 points; removing both edge families costs 5.0; withholding biography text (index only) drops to 68.0% (−4.0); withholding biography text only from the answer model drops to 70.8% (−1.2); removing references to unsearched observations drops to 70.2% (−1.8); removing caption text drops to 61.8% (−10.2); removing visual frames drops to 70.4% (−1.6).
  • Ablations on MM-Lifelong: Appending descriptions to captions costs 3.25 points on Test@Week and 5.08 on Test@Day; withholding biography text with the index retained costs 6.75 and 2.33 respectively. Neither variant recovers full performance.
  • Confidence intervals: Intervals exclude zero for EgoLifeQA, Ego-R1-Bench, and Test@Week, but include zero for Test@Day.
  • Name ambiguity is measurable in the corpus: 89,888 same-frame pairs of proven-different objects carry the same described name, occurring in 5,804 of the 6,266 clips (92.6%). Conversely, under the authors' association, 77.2% of the 22,401 object entities seen more than once are described under two or more distinct names.

Methodology in Plain English

The researchers split a recording into short clips (the EgoLife week is 6,266 thirty-second clips at 1408×1408, sampled at 10 fps). An open-vocabulary detector (YOLOE-26x-seg in prompt-free mode) finds objects on sampled frames, BoT-SORT links detections into within-clip tracks, and tracks that the tracker split are regrouped by CLIP and DINOv2 appearance similarity, never grouping tracks seen in the same frame. A group visible for at least two seconds becomes an "observation," keeping representative crops, its time span, and its source episode.

Each observation is described by a single call to a vision-language model (Qwen3.5-35B) that sees four crops, scene frames with the subject boxed, and the recording's own caption and transcript within ±40 seconds. The prompt tells the model to establish the subject visually and use textual context only when it concerns that subject.

Association is the central step and is deliberately conservative. Observations are processed in temporal order. Each new observation is embedded (Qwen3-VL-Embedding-8B) and compared against each existing entity's 10 most recent assigned observations, used as references. A match requires three conditions: the new observation must be at least as similar as a floor threshold to every reference, strongly similar to at least one, and show no visual separation from any reference. Separation is measured geometrically as bounding-box IoU on shared frames: two observations count as separate when they share at least three frames and their boxes overlap with IoU below 0.5 in at least two thirds of them. Thresholds are set per recording (0.60/0.75 on EgoLife, 0.75/0.85 on the gameplay stream, whose repeated assets make objects of one kind score alike). If no candidate qualifies, the observation starts a new entity.

The memory is a graph with persistent entity nodes, observation nodes, episode nodes at multiple time scales (30-second, 3-minute, 10-minute, and 1-hour on EgoLife), and source-clip nodes. Edge weights are episode context 1.0, visual provenance 0.8, and same-instance 0.5. Retrieval indexes observation descriptions and episode captions together and propagates relevance through Personalized PageRank (damping 0.85), ranking nodes by PageRank score times query-text cosine similarity, restricted to records preceding a timestamped query. A controller summarizes each entity's retrieved observations as a biography excerpt in time order and lists unselected appearances by times and counts, giving the controller concrete targets for the next search round. Retrieval allows at most five rounds and 64 frames; week-scale retrieval uses 16 units per round (at most six observations, at most three from one entity), while MultiHop-EgoQA uses three units per round with at most one observation.

Why This Matters

Impact on research: The paper reframes long-video memory as an identity-resolution problem rather than only a compression or retrieval-ranking problem. It provides a measurable demonstration that language-derived entity links conflate distinct physical instances and fragment single instances across names, and it shows that resolving identity in the memory improves both answer accuracy and access to supporting moments under matched controller, answer model, and retrieval limits.

Real-world applications:

  • Personal and wearable assistants that must answer questions like whether a specific item was used, moved, or put away, where many similar objects (mugs, tools, cables) share descriptions.
  • Robotics and embodied agents that need persistent object histories across sessions for manipulation and household tasks, where continuous tracking is not available.
  • Video surveillance, industrial, or safety review, where the same equipment or vehicle must be followed across hours of footage and across camera cuts or recording sessions.
  • Long-form media and sports or gameplay archives, where recurring characters, assets, or players must be followed across a lengthy stream; the paper evaluates a 23.6-hour gameplay stream for exactly this recurring-entity setting.

Industry relevance: The framework builds on off-the-shelf components (open-vocabulary detection, tracking, embedding models, a 35B vision-language model, PageRank) and writes the memory once, offline, to serve all later questions. The reported scale — 308,244 tracked observations, 77,716 entities, 27,446 entities with two or more observations, 15,365 spanning more than one day — indicates the approach is intended for production-scale archives. It also integrates with existing retrieval controllers rather than replacing them: GEB and MAGIC-Video share episodic captions, topic and event summaries, and retrieval limits.

Future Directions

  • Reducing identity fragmentation: Conservative association can leave one physical instance under multiple identifiers; the paper notes a reading-time note that distinguishes pairs with visual evidence of separation from unresolved ones, leaving automated merging of split entities open.
  • Closing the whole-clip gap: On MultiHop-EgoQA, the answer model reading 60 frames of the whole clip without memory scores higher (3.64) than GEB (3.11), and GEB's complete coverage (0.521) remains below WorldMM's (0.526), so retrieval-side evidence selection is not fully solved.
  • Extending naming and identity beyond named casts: The naming stage runs on the EgoLife recording, whose seven participants are named; the Test@Day livestream has no cast, so its entities carry only names written in their descriptions.
  • Broadening benchmark coverage and stabilizing weaker splits: The Test@Day confidence interval includes zero, and the paper tests only four benchmarks, three of them centered on the same EgoLife week.

Target Audience

Researchers and engineers working on long-video understanding, video question answering, agentic retrieval, and multimodal memory systems. It is also relevant to practitioners building retrieval or knowledge-graph pipelines over video archives, and to readers interested in entity resolution and identity grounding as a distinct problem from captioning or description retrieval. The paper is best suited to readers comfortable with graph retrieval and vision-language model pipelines; the ablation and evidence-access analyses are accessible to a broader technical audience.

Authors’ abstract

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

Read the original paper