Skip to content
AI.info

Research

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

Overview Research area: Computer vision and efficient vision-language modeling — specifically visual-token compression and memory organization for streaming and long-video question answering. Technica

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
arXiv
2609.32182
Published
2026-09-26
Authors
Zihan Chen, Xuejian Rong, Xiaojuan Wang, Boqing Gong, Adi Zicher, Yael Pritch, Nikhil Karnad

AI summary

Overview

Research area: Computer vision and efficient vision-language modeling — specifically visual-token compression and memory organization for streaming and long-video question answering.

Technical level: Advanced.

Scope: The paper introduces KeyRec, a training-free framework that builds a bounded, query-agnostic visual memory (a recent cache plus a structured key-event bank) and reads from it with a query-adaptive, fixed token budget, evaluated across four benchmarks and three VLM backbones.

What This Paper Is About

Vision-language models represent each video frame with dozens or hundreds of visual tokens, so the visual context grows linearly with video length while attention cost grows quadratically, making long streams and long recordings expensive or infeasible to process. Existing training-free token-selection methods drop redundant tokens but tend to fragment evolving events, discard the context that makes an event coherent, and treat detailed recent observations and long-range history under one uniform retention policy. KeyRec instead treats compression as memory organization: it separates a dense recent cache from a bounded bank of temporally localized events, then lets the question itself decide how much of a fixed readout budget goes to recent versus historical evidence.

Key Contributions

  1. A bounded visual-memory framework. KeyRec decouples query-agnostic memory construction from query-adaptive, fixed-budget readout, formalized with a persistent storage budget S and a per-query readout budget B.

  2. A semantic-temporal event bank. History is organized into temporally localized events maintained by an online add–merge–evict update: candidates are proposed by novelty relative to stored event routes, merged only when both a similarity threshold and a temporal-locality constraint hold, and evicted by a retention score that balances event utility against temporal coverage.

  3. A text-only router for budget allocation. The same frozen VLM reads only the question and a textual description of the two memory sources, choosing one of five readout states (AllRecent, HeavyRecent, Balanced, HeavyEvent, AllEvent) that split the budget between recent and event memory and choose between relevance-oriented retrieval and uniform temporal coverage.

  4. Support for native one-vision models. Because KeyRec operates on model-facing visual embeddings, it transfers to the encoder- and projector-free NEO-ov architecture with only structure-preserving positional bookkeeping (frame index and spatial location), without changing the memory algorithm or reconstructing dense visual grids.

Main Findings

  • Best compressed result in 13 of 15 settings at a 10% budget. Using approximately 10% of the dense decoder-facing visual-token budget, KeyRec achieves the best compressed performance in 13 of 15 settings across four benchmarks and three VLM backbones.

  • Strong streaming gains. KeyRec achieves the best compressed performance in eight of nine streaming settings and outperforms the strongest compressed baseline by 2.21–18.37 points across OVO-Bench Realtime and StreamingBench. It matches or exceeds uncompressed Vanilla in all nine streaming settings, with gains of up to 10.01 points.

  • Vanilla is not an oracle. The paper argues that redundant streaming prefixes can dilute decisive recent evidence, so a bounded, structured memory can beat dense inference rather than merely approximating it.

  • Honest long-video trade-off. KeyRec obtains the best compressed result in five of six long-video settings and consistently outperforms other compressed methods on LongVideoBench, but remains 0.30–3.37 points below Vanilla on those settings — the fixed 10% readout preserves much of long-video capability without retaining all fine-grained evidence.

  • Consistent advantage on NEO-ov 2B. KeyRec achieves the best compressed performance in every NEO-ov 2B setting; in the streaming OVO-Bench Realtime setting it scores 62.78 versus 44.41 for CausalMem, whose larger retained-token buffers grow with stream length.

  • Scaling behavior differs by budget. On StreamingBench, KeyRec improves rapidly at small budgets, peaks at a 5% readout, and then declines as visual context grows; the peak configurations also outperform uncompressed Vanilla.

  • Architecture-dependent frame scaling. On LongVideoBench videos between 900 and 3,600 seconds at a fixed 10% ratio, dense Vanilla becomes infeasible — Gemma runs out of memory at 512 frames and NEO-ov at 128 frames or more. On NEO-ov, STC-Pruner and CausalMem degrade substantially while KeyRec remains stable across 32–512 frames, staying above 44% accuracy at 512 frames versus below 25% for both baselines.

  • The comparison controls for storage. KeyRec selects each readout from a persistent memory roughly twice as large: at approximately matched persistent storage, KeyRec is compared at a 5% readout against baselines at 10%, and at 10% against baselines at 20%, indicating the gain comes from organization and adaptive readout rather than simply storing more tokens.

  • Allocation policy matters, and no fixed policy wins. On Gemma 4 E4B, OVO-Bench Backward benefits most from event memory, OVO-Bench Realtime degrades sharply under an event-only readout, and LongVideoBench performs best with a mixture — the router matches the event-only policy on backward reasoning while outperforming fixed policies on real-time and long-video understanding.

  • Modest latency cost, better scaling. On LongVideoBench with Gemma E4B, KeyRec's end-to-end latency overhead relative to StreamingTOM-CTR decreases from approximately 19% at 64 frames to 10% at 512 frames. At 512 frames, dense Vanilla runs out of memory; StreamingTOM-CTR takes 39.12 s and 16.57 GiB, and KeyRec takes 43.08 s and 18.93 GiB. Vanilla already reaches 39.23 GiB at 256 frames.

  • Full-prefix 1-FPS evaluation. With a fixed decoder-facing budget of B = 256 × 70 × 0.1 = 1,792 visual tokens independent of prefix length (OVO-Bench mean duration approximately 235 s, StreamingBench approximately 264 s), KeyRec achieves higher overall performance than CausalMem on both benchmarks across E2B and E4B, with clear gains on real-time understanding; on OVO-Bench Backward it improves E2B (48.61 versus 46.70) and stays close on E4B (56.15 versus 56.58).

  • Serialization matters. Presenting selected tokens as neutral numbered memory segments (Default) generally beats both flat concatenation (Naive) and semantically labeled segments (Detailed). On Gemma 4 E4B, Default scores 54.45 / 64.81 / 45.77 on OVO-Bench Backward, OVO-Bench Realtime, and LongVideoBench, versus Naive 56.40 / 63.72 / 41.21 and Detailed 53.10 / 63.78 / 44.43; on NEO-ov 2B, Default scores 41.90 / 62.78 / 49.29 versus Naive 38.78 / 62.01 / 48.84 and Detailed 39.86 / 61.90 / 48.99. The paper attributes this to preserving memory organization while avoiding semantic priors such as the word "key" biasing how evidence is weighed.

Methodology in Plain English

KeyRec works in two phases. In the first, which does not look at any question, each incoming frame updates two things. A recent cache simply keeps the most recent N_C visual tokens, giving fine-grained detail about what just happened. A key event memory stores at most K events, each with a compact route embedding, at most m visual tokens, their timestamps, and a utility value. Each frame proposes a candidate event built from its m most novel tokens, where novelty is measured as one minus the highest cosine similarity to any stored event route — so novelty means "not yet represented in history," not "salient within this frame." The candidate is either inserted as a new event or merged into the most similar stored route, but merging additionally requires that the matched event be temporally adjacent (controlled by thresholds γ for similarity and δ for temporal gap), which prevents visually similar occurrences at distant times from being collapsed together. Merged events keep a normalized moving average of their route, re-sort and uniformly subsample their tokens and timestamps to stay within m, and take the maximum of the old and new utility. When the bank exceeds K events, the event with the lowest retention score — utility plus a term rewarding center times far from other events — is evicted.

At query time, a text-only pass of the same frozen VLM classifies the question into one of five states. The state sets the fraction π_q of the readout budget given to events; the number of events and recent tokens follow directly from B and m. HeavyRecent and Balanced select events by cosine similarity between event routes and the mean-pooled question embedding; HeavyEvent and AllEvent select events by uniformly spreading across center times for temporal coverage; AllRecent takes no events. The selected tokens are restored to temporal order and inserted as separate labeled segments at their cached model-facing embeddings, so no historical frame is ever reprocessed.

Why This Matters

The paper reframes visual-token compression from "which tokens are important" to "how should visual evidence be organized and presented." It shows that a bounded memory plus an adaptive readout can beat uncompressed inference on streaming benchmarks, and it opens the underexplored question of how compression methods transfer to encoder-free, one-vision architectures such as NEO-ov 2B, where ViT-specific attention signals and positional mechanisms are unavailable.

Real-world applications:

  • Always-on assistants on phones, cameras, or smart glasses that must answer questions about a live feed without knowing how long it will run or when a question will arrive.
  • Real-time monitoring of long-running processes, where a system must distinguish "what is happening right now" from "what happened earlier" under a fixed compute budget.
  • Search and QA over long archives such as full-length recordings, where the same video is queried repeatedly and should not be reprocessed for each question.
  • Edge and on-device deployment, where persistent storage and per-query decoding cost must stay bounded independently of stream duration.

Industry relevance: the results matter to anyone serving video-language models at scale, because dense prefilling cost grows quadratically and dense memory grows linearly. KeyRec offers a training-free drop-in that keeps a persistent memory at roughly 20% of dense visual input while exposing only 10% to the decoder per query, delivering higher accuracy than the strongest compressed baseline at a latency overhead that shrinks from approximately 19% to 10% as videos grow from 64 to 512 frames.

Future Directions

  • Adapting visual memory to encoder-free VLMs. The paper explicitly names this an underexplored and consequential direction, noting that as architectures such as NEO-ov and Gemma 4 12B emerge, architecture-specific positional mechanisms and visual-token interfaces differ from conventional encoder–projector VLMs, and transferable methods may change substantially in effectiveness.

  • Event granularity and storage allocation. The paper includes a sensitivity study that varies per-event token size m while adjusting event-bank capacity K inversely, holding K·m = 256 on OVO-Bench, but the reported content is truncated before the results — the outcome of this study is not reported in the available text.

  • Learning or refining the readout policy. The current router is a text-only classifier over five discrete states with fixed retrieval strategies; how finer-grained or learned allocation would behave is not explored here.

  • Multimodal inputs beyond vision. All evaluations exclude subtitles and audio to isolate visual-memory compression, so how the bounded memory should incorporate speech or text streams is left open.

Target Audience

Researchers and engineers working on efficient vision-language models, long-video and streaming video understanding, and visual-token compression; practitioners deploying VLM-based video systems under memory or latency constraints; and anyone interested in how compression techniques must be redesigned for emerging encoder-free, one-vision architectures.

Authors’ abstract

Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.

Read the original paper