Skip to content
AI.info

Research

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants Overview Research area: Computer vision / streaming video understanding, specifically egocentric (firs

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
arXiv
2609.37559
Published
2026-09-29
Authors
Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li, Zhimian Wen, Mingda Li, Shule Lu, Zhicheng Wang, Yuhan Guo, Xin Jin, Wenjun Zeng

AI summary

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

Overview

Research area: Computer vision / streaming video understanding, specifically egocentric (first-person) video assistants and long-term memory for multimodal models.

Technical level: Advanced. The paper assumes familiarity with streaming video benchmarks, KV-cache memory, visual token compression, and LLM-as-judge evaluation protocols.

Scope in one sentence: APM-Bench is a new benchmark of 549 sessions, 104 trajectories, and 2,719 human-refined candidates that measures whether streaming video assistants can store, selectively retrieve, and reuse persistent memory across temporally separated interactions while balancing task utility, response latency, and storage cost.

What This Paper Is About

Existing streaming video benchmarks evaluate models on a single continuous video or a short clip, but real-world assistant use is intermittent: a user may turn off smart glasses and resume hours or days later. APM-Bench reformulates streaming interaction as multi-session "life trajectories" of related activities, so that a model must retain visual evidence from completed sessions and reuse it later. The paper's goal is to measure whether such persistent memory is simultaneously useful, cheap to store, fast to access, and honest about what evidence it no longer has.

Key Contributions

  1. APM-Bench itself. A benchmark containing 549 sessions, 104 trajectories, and 2,719 candidates spanning objective (multiple-choice) and open-ended questions, averaging 69 minutes of video per trajectory with individual sessions averaging 13 minutes. It is built from EgoLife (multi-day recordings, transcripts, timestamped captions) and HD-EPIC (fine-grained action and object annotations for structured kitchen procedures), with an inter-annotator Cohen's kappa of 0.868 on 300 sampled questions.

  2. A three-capability, 12-task evaluation structure. Cross-session Understanding (Episodic Recall, Entity State Tracking, Temporal Reasoning), Real-time Perception (Action Recognition, Counting, OCR, Spatial Understanding), and Adaptive Response (Evidence-Ready Answering, Registered-Condition Response, Memory-Grounded Proactive Assistance, Proactive Reminder, Task Progress Guidance), spanning six streaming temporal formulations and both intra-session and inter-session evidence locations.

  3. An Evidence Availability-Aware evaluation set. A dedicated set of 260 Cross-session Understanding questions, split into 130 where all required evidence lies within the two most recent completed sessions and 130 where it does not, answered as 5-way MCQA with a fifth option for insufficient available evidence.

  4. A systematic comparison of memory representations and systems. Evaluation of general video models under three memory protocols (no memory, raw video as memory, text summary as memory) and eight specialized streaming memory systems across five memory representations (KV Cache, Visual Tokens/Features, Event Tree, Parametric Memory, Reasoning Thoughts), reported alongside time-to-first-token and storage cost per video hour.

Main Findings

  • Raw video is the strongest but heaviest memory. Raw Video as Memory gives the best Cross-session Understanding for most general models: Gemini 3.6 Flash reaches 69.37 and Seed-2.0-Lite 54.47, against 27.21 for the no-memory SimpleStream baseline. The paper reports this comes with GiB-scale storage (Seed-2.0-Lite: 3.01 GiB per video hour, Table 3) and high query-time overhead, and the paper states raw-video memory "often weakens real-time perception."

  • Text summaries cut storage by orders of magnitude but lose visual detail. Text Summary as Memory stores 7.75 KiB (Seed-2.0-Lite), 3.09 KiB (Gemini 3.6 Flash), and 6.84 KiB (Qwen3.8-27B) per video hour, but Cross-session Understanding drops to 45.34, 47.50, and 41.05 respectively for those same models.

  • Event-structured memory leads among specialized systems. StreamForest posts the highest Cross-session Understanding among the eight specialized systems at 37.79, and OASIS is the only specialized system reported as supporting both streaming input and persistent memory, with the highest overall score among specialized systems at 41.46. Parametric memory (Video-Salmon-S, overall 18.72) and reasoning-thought memory (VST, overall 17.54) trail.

  • A clear utility–latency–storage trade-off exists. The paper states that no evaluated method simultaneously achieves strong task performance, low latency, and low storage. Reported per-video-hour storage ranges from 11.64 KiB (VST) and 0.33 GiB (FLUXMem) up to 20.84 GiB (ReKV), while time-to-first-token ranges from 0.56 s (SimpleStream) to 36.15 s (StreamForest) and 32.44 s (OASIS).

  • More history can hurt real-time perception. Replacing raw-video history with text summaries improves Real-time Perception for most general video models (Gemini 3.6 Flash: 64.20 to 65.98; Seed-2.0-Lite: 50.75 to 52.44), which the authors read as evidence that excessive or irrelevant history can interfere with current-scene understanding.

  • Adaptive Response is the hardest capability. Even with Raw Video as Memory, the strongest general model scores 47.05 on Adaptive Response. Tasks with an explicit registration do better than fully autonomous ones: the paper reports RCR and PRM substantially outperform MPA and TPG, where no explicit trigger specifies what historical experience is relevant. Specialized systems are very weak here (StreamForest 6.48, Flash-VStream 11.41, VST 15.24).

  • Recall ability and evidence-availability awareness are decoupled. On the Evidence Availability-Aware set, Gemini 3.6 Flash scores 85.38 when the evidence is accessible but 59.23 at detecting unavailable evidence; SimpleStream shows the opposite pattern, scoring 3.85 when evidence is available and 95.38 at flagging it as unavailable (where it overwhelmingly selects the insufficient-evidence option).

Methodology in Plain English

The researchers took existing egocentric datasets (EgoLife and HD-EPIC) and reorganized them into "trajectories" — chains of sessions that share related activities and preserve real wall-clock ordering and gaps. Each session carries fine-grained annotations of evidence intervals, query times, and reference proactive responses. Candidate questions and probes were generated from session videos and then filtered and refined through choice-blind review, timestamp filtering, video-agent review, and human verification.

At evaluation time, models see the current session's causal video prefix up to a probe timestamp plus whatever persistent memory they are allowed to hold from prior, completed sessions. Cross-session Understanding and Real-time Perception are scored as 4-way multiple choice (887 and 716 runtime instances respectively). Adaptive Response is scored over 3,768 probes, each treated as an independent runtime instance where the model must decide SILENT versus INTERVENE; a Gated LLM-Judge first checks the decision with a deterministic gate (and the MCQA option for positive Evidence-Ready Answering probes), then has DeepSeek-V4-Flash rate the rationale on a 1–5 scale, scaled by a factor of 20 to a 0–100 task score. Storage cost is measured as peak prior-session memory footprint normalized by cumulative prior-session hours. Storage is estimated from the memory state maintained during inference for systems that do not export persistent state. All experiments ran on H20 GPUs, with video input at 1 FPS; proprietary and Qwen-series

Authors’ abstract

To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.

Read the original paper