Research
MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
Overview Research area: Computer vision / egocentric video question answering, specifically long-term video memory systems and retrieval-augmented generation over wearable-camera footage. Technical le

- arXiv
- 2609.40195
- Published
- 2026-09-30
- Authors
- Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar, Raffay Hamid, Aidong Zhang, Xin Luna Dong
AI summary
Overview
Research area: Computer vision / egocentric video question answering, specifically long-term video memory systems and retrieval-augmented generation over wearable-camera footage.
Technical level: Advanced. The paper combines agentic retrieval tool design, reinforcement learning with group-relative policy optimization (GRPO), and an information-theoretic decomposition of answer uncertainty.
Scope: The paper introduces MemLife, a multimodal memory system that compresses long-term egocentric video into entity-grounded, first-person text episodes read by a time-indexed agentic reader, and MemOpt, an RL framework that trains only the memory writer against a Faithful–Informative–Retrievable Memories (FIRM) reward.
What This Paper Is About
Personal AI assistants that record daily life through wearables would need to answer arbitrary questions about months or years of footage, but reprocessing raw video for every query is computationally prohibitive. Compacting video into text memory is a scalable alternative, yet existing systems fail in two ways: the written memory drops key evidence, or the retriever cannot find the right entry as the search space grows. The paper's goal is to build a memory system that remembers only what is worth remembering, faithfully, in a form that is easy to recall.
Key Contributions
-
MemLife, a multimodal memory system that converts egocentric video into time- and entity-anchored, first-person text episodes and reasons over them with a versatile agentic reader. The reader combines agentic semantic search with time-scoped memory fetching, presents retrieved episodes in chronological order, and can optionally invoke video-retrieval tools to sample raw frames (MemLife-V).
-
MemOpt, a reinforcement learning framework whose FIRM reward optimizes the memory writer with multi-granular feedback on faithfulness, informativeness, and retrievability. Unlike prior work that applies RL to final answer quality, MemOpt applies RL only to memory writing and keeps the agentic reader fixed, decoupling writer post-training from reader execution noise.
-
Strong empirical results: MemLife + MemOpt outperforms the strongest prior baseline by 4.0–17.0% while reducing memory size by up to 31× (Appendix E). MemLife alone improves over the strongest training-free baseline by 4.6–12.0% across four long-horizon benchmarks, and MemOpt adds 2.7–5.0%.
-
Demonstrated transfer: the trained writer consistently improves accuracy when plugged into other memory systems (e.g., EgoRAG) and backbones, and on out-of-domain video and question types.
Main Findings
-
MemLife beats training-free baselines without training or query-time video access. On SuperMemory-VQA, MemLife reaches 56.50 accuracy versus 49.28 for EgoRAG, the strongest training-free baseline; on EgoLifeQA, 52.80 versus 48.20 (EgoRAG); on SuperMemory-LVQA, 56.18 versus 49.28 (Video-RAG); on EgoLife-EQA, 52.00 versus 40.00 (VideoARM).
-
Video access changes results only modestly. MemLife-V (source-video access enabled) produces 57.78 on SuperMemory-VQA, 53.20 on EgoLifeQA, 57.95 on SuperMemory-LVQA, and 50.00 on EgoLife-EQA, showing the gains do not depend on revisiting original recordings.
-
MemOpt improves MemLife on every benchmark and beats every trained baseline. MemLife + MemOpt scores 60.35 (SuperMemory-VQA), 56.60 (EgoLifeQA), 58.91 (SuperMemory-LVQA), and 57.00 (EgoLife-EQA). MemLife-V + MemOpt scores 60.83, 57.60, 61.96, and 57.00 respectively. All other accuracy figures in the paper are below the Oracle Context reference of 67.58, 66.20, 67.58, and 61.00.
-
Trained baselines trail. With writer training, the next-best trained system on SuperMemory-VQA is M3-Agent at 53.93; on SuperMemory-LVQA, M3-Agent at 54.90; on EgoLifeQA, TaskMem at 46.20; on EgoLife-EQA, VST at 36.00.
-
Recall changes are mixed. MemOpt improves accuracy on all benchmarks but recall changes are mixed, indicating the accuracy gains also reflect more answer-useful memory content rather than only retrieving annotated evidence.
-
MemLife improves both retrieval and memory content. Against EgoRAG on EgoLifeQA question categories, MemLife improves recall in every category and raises oracle accuracy overall and in most categories. On RelationMap, EgoRAG matches the optimized MemLife system in standard accuracy despite lower recall and oracle accuracy.
-
MemOpt learns concrete memory corrections. Examples show the untrained writer mistaking lentils for corn and using a vague pronoun while omitting the relevant food; MemOpt corrects the object and identifies the person and the baked chicken needed to answer the question, removing unsupported details while making answer-relevant entities explicit.
-
Gains hold across writer and reader backbones. With a Qwen3.5-9B writer, zero-shot accuracy is 56.50 (SuperMemory-VQA) and 52.80 (EgoLifeQA) under the Qwen3.5-9B reader, rising to 60.35 and 56.60 with MemOpt; under a Qwen3.6-27B reader, 66.93 and 59.60 rise to 67.90 and 60.20. With a Qwen3.6-27B writer, zero-shot figures are 57.14/55.20 (9B reader) and 66.77/59.80 (27B reader), rising to 60.67/56.60 and 68.06/60.60. Gains become smaller with the stronger reader.
-
MemOpt transfers across memory systems. Training improves the native EgoRAG pipeline on both benchmarks; using MemLife to read optimized EgoRAG memories provides further gains, and the highest performance comes from the trained MemLife writer.
-
Writer components are additive. Progressively adding entity grounding and first-person narration to multimodal fusion raises SuperMemory-VQA zero-shot accuracy from 52.33 to 54.09 to 56.50, and MemOpt accuracy from 56.98 to 59.87 to 60.35; first-person narration consistently raises EgoLifeQA recall (from 45.60 to 48.40 zero-shot, 47.60 to 50.60 with MemOpt).
-
Reader components are additive. Adding time anchoring and chronological ordering to agentic reasoning raises MemLife zero-shot EgoLifeQA accuracy from 50.40 to 50.60 to 52.80 and MemOpt EgoLifeQA accuracy from 51.80 to 54.00 to 56.60, with the complete reader performing best overall.
-
The full reward matters. Teacher imitation with Qwen3.6-27B gives only modest gains (57.78 on SuperMemory-VQA, 53.00 on EgoLifeQA); retrievability alone raises recall but not consistently accuracy; adding informativeness strongly benefits SuperMemory-VQA (60.03) but not EgoLifeQA (51.40); adding faithfulness produces the best accuracy on both benchmarks (60.35 and 56.60).
-
Token-level faithfulness plus multiplicative aggregation wins. Sequence-level/additive gives 58.75 (SuperMemory-VQA) and 55.00 (EgoLifeQA); token-level/additive gives 58.91 and 56.00; token-level/multiplicative gives 60.35 and 56.60.
-
Context conditioning did not help. Conditioning the writer on textual or multimodal context from preceding segments improved neither variant's aggregate accuracy (Appendix I).
Methodology in Plain English
The system has two parts: a writer and a reader.
The writer converts each video segment — a sampled frame plus its aligned audio transcript and start/end timestamps — into a short text description stored with its timestamp. Three design choices guide it. Multimodal fusion means the writer reads frames and transcript together so evidence from either stream can survive. Entity grounding means names mentioned in speech are used to refer to the people and objects seen in the clip. First-person narration means descriptions are written as "I" statements, matching how a wearer asks about their own life. Each clip is described independently, so computation and storage stay linear in the recorded history and no errors propagate between segments.
The reader is an agent that takes the question and its timestamp and repeatedly picks from five actions: Rewrite (turn the question and time into a search query and/or a time interval), SearchMemory (similarity search returning k entries inside an interval), FetchMemory (return all text entries in an interval), FetchVideo (return sampled frames in an interval, enabled only when source video is available), and Answer. Retrieved episodes are presented in chronological order.
MemOpt trains the writer while the reader stays fixed. The authors first decompose answer uncertainty into two losses: the memory writing loss (information lost when video is mapped to memory) and the memory recall loss (information lost when the reader retrieves from memory). This motivates two of the three reward terms. Faithfulness is assessed token by token: the same frozen model checks the memory against its source segment, reproducing supported content and minimally correcting unsupported spans, and the reward is 1 minus the gap between the evaluator's top token probability and the probability of the writer's actual token. Informativeness checks whether the memory entails a source-grounded key fact, averaged over the questions that segment can answer. Retrievability checks whether the memory is returned for those questions, using reader actions cached at the start of each epoch and replayed rather than rerun. The three rewards are combined multiplicatively at the token level, so a token earns high credit only if it is faithful and its memory is both informative and retrievable. Advantages are computed by averaging within each candidate before group normalization, preventing longer memories from dominating, then training follows GRPO.
Training and evaluation data: MemOpt is trained on SuperMemory-VQA subjects S1–S6 (3,425 questions), validated on S7–S8 (723), and tested on S9–S10 (623), with 82 questions excluded because their annotated evidence refers to unavailable source videos, leaving 4,771 questions total. EgoLifeQA (500 multiple-choice questions over a continuous seven-day recording), SuperMemory-LVQA (all ten SuperMemory-VQA histories concatenated into one store, retaining the 623 test questions), and EgoLife-EQA (100 questions about recurring events, with two annotators independently verifying evidence at 92.1% agreement) are test-only. Qwen3.5-9B is the backbone for both writer and reader across systems, with Qwen3.6-27B used in the generalizability study.
Why This Matters
Impact on research. The paper reframes long-term video QA as a memory-writing problem rather than an answer-generation problem, showing that training only the writer — without stronger-model supervision and without rerunning a multi-turn reader during training — improves downstream accuracy. Its information-theoretic decomposition separates writing loss from recall loss, giving a principled justification for two reward terms, and its demonstration that optimized memories transfer to other memory systems (EgoRAG) and backbones suggests reusable memory writers rather than system-specific pipelines.
Real-world applications:
- Personal memory assistants on smart glasses or GoPro-style wearables that answer questions such as "Where did I put my passport?" or "Did I add the milk before or after the eggs?"
- Life-logging and health/behavior review, where a user asks which days an activity occurred or which of two activities happened more often.
- Enterprise or field-work recording, where a worker needs to recall procedural details captured across many sessions.
- Privacy-conscious deployments, since MemLife stores low-resolution redacted video (due to storage and privacy concerns) and samples limited frames only when visual detail is required.
Industry relevance. The 31× memory-size reduction matters directly for storage and serving cost on always-on wearable devices, and the latency budget for sifting through long, similar memories is described as tight. Training only the writer, with cached reader actions for retrievability supervision, is a cheaper post-training recipe than end-to-end task-reward RL, which the paper characterizes as noisy and computationally prohibitive.
Future Directions
- Closing the gap to Oracle Context. Oracle Context reaches 67.58, 66.20, 67.58, and 61.00 across the four benchmarks, well above the best reported MemLife/MemOpt numbers, leaving substantial headroom.
- Understanding the accuracy–recall inconsistency. On SuperMemory-LVQA methods maintain accuracy close to SuperMemory-VQA despite lower annotated recall, and the authors hypothesize cross-subject interactions supply useful context outside annotated evidence — an open question about what "correct retrieval" means in multi-subject stores.
- Better memory for relational questions. On RelationMap, EgoRAG matches the optimized MemLife system in standard accuracy despite lower recall and oracle accuracy, suggesting evidence-aligned memories do not reliably preserve relations.
- Adapting the reader to compensate less. MemOpt gains shrink with a stronger Qwen3.6-27B reader, raising the question of how writer optimization and reader capacity should be balanced.
- Extending supervision beyond one dataset. MemOpt is trained only on SuperMemory-VQA subjects S1–S6, so scaling and diversifying training data remains unexplored.
Target Audience
Researchers and engineers working on egocentric video understanding, long-term video question answering, memory-augmented LLM agents, and retrieval-augmented generation over multimodal streams. It is also relevant to practitioners building wearable or personal-assistant products who care about memory footprint, inference cost, and privacy-preserving video handling. Readers need familiarity with retrieval-augmented generation, agentic tool-calling, and reinforcement learning with group-relative advantage estimation.
Authors’ abstract
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text representations, but often fail on practical benchmarks: either the memory does not preserve key evidence, or the retriever fails to locate relevant entries due to retrieval competition in growing search spaces. To address these challenges, we introduce MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. Without training or query-time video access, MemLife improves over the strongest training-free baseline by 4.6--12.0% across four long-horizon benchmarks. To further improve memory quality, we propose MemOpt, a reinforcement learning framework that optimizes the memory writer to produce faithful, informative, and retrievable memories. MemOpt consistently improves MemLife by 2.7--5.0% across different video and question distributions, with gains that generalize across writer and reader backbones and memory systems.