Research
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
Overview Research area: Computer Vision / multimodal video-language models — specifically streaming (online) video large language models that perceive a live stream, build memory, and decide when to a

- arXiv
- 2610.01762
- Published
- 2026-10-01
- Authors
- Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen, Yansong Shi, Junhao Zhou, Yifei Li, Jun Zhang, Chuanyu Qin, Chenxu Yang, Xinlei Yu, Kun Ouyang, Yuchen Shao, Qianshan Wei, Changhai Zhou, Jun Gao, Jiaqi Wang, Limin Wang
AI summary
Overview
- Research area: Computer Vision / multimodal video-language models — specifically streaming (online) video large language models that perceive a live stream, build memory, and decide when to answer.
- Technical level: Advanced. The paper assumes familiarity with video LLMs, autoregressive decoding with control tokens, sliding context windows, parameter-efficient fine-tuning, and benchmark suites for online video understanding.
- Scope: The paper introduces OneStreamer, a 4B-parameter streaming video LLM that unifies query-independent memory writing and proactive task response through a shared generation process, together with a data synthesis pipeline and the OneStreamer-1M training corpus.
What This Paper Is About
A streaming video model sees video arrive as an open-ended stream and does not know in advance which moments will matter to a future question. If it keeps all past visual tokens, those old tokens compete with current frames for limited context capacity and can weaken real-time perception; if it keeps only a recent window, distant evidence is lost. OneStreamer's goal is to turn observed video into reusable, time-grounded factual records that survive after their source frames leave the visual window, while also learning when enough evidence has arrived to produce a user-visible response.
Key Contributions
- Proactive Hierarchical Caption Memory (PHCM). The model generates two kinds of time-aligned textual records from already-observed content: dense local-detail captions introduced by
</Observe>and sparser semantic summaries of completed events or segments introduced by</Summary>. These records are retained as text after their source frames leave the Recent-N visual window, so the visual window itself is never changed or revisited. - Proactive State Transition Learning (PSTL). A selective supervision scheme for control-state tokens that preserves supervision at all output anchors and selects representative state-change and state-persistence tokens, while masking unselected state tokens only from the state loss.
- A streaming data synthesis pipeline with a caption branch (video curation, multi-granularity annotation with Gemini and Seed, verification, causal streaming sequence construction) and a QA branch (task-directed QA design, coarse evidence interval localization, fine-grained response-time calibration).
- OneStreamer-1M, a broad-coverage streaming video interaction dataset whose listed records total 1,160,004, combining newly synthesized OneStreamer-Core data with curated open-source data (from Streamo, JoyAI-VL-Interaction, MMDuet2, ViSpeak, StreamForest, VideoChat-OL, Seeker, and VST), decontaminated at the source-video level against every evaluated benchmark.
Main Findings
- Best results on all eight evaluated benchmarks. The 4B OneStreamer scores 72.1 on OVOBench, 86.9 on StreamingBench Real-Time, 66.8 on OVBench, and 71.3 on ODVBench among perception and memory benchmarks; and 48.7 on ProactiveVQA, 36.6 on OmniMMI, 41.6 on OVO-Timing, and 2.87 on ViSpeak among proactive-response benchmarks.
- Size-matched and larger-baseline gains. Against the 4B Qwen3-VL base model it improves OVOBench, StreamingBench Real-Time, OVBench, and ODVBench by 13.3, 5.1, 11.4, and 13.7 points, and ProactiveVQA, OmniMMI, and OVO-Timing by 14.4, 7.2, and 12.2 points. It surpasses the 11B MOSS-VL-Realtime by 1.9, 4.0, 13.1, and 7.4 points on the four perception/memory benchmarks and by 1.5, 3.9, and 3.1 points on the first three proactive benchmarks; on ViSpeak it scores 2.87 versus 2.41 for Qwen3-VL and 2.48 for MOSS-VL-Realtime.
- Aggregate improvement reported in Figure 1. The paper reports an average relative improvement of 25.0% over the Qwen3-VL baseline and 6.1% over the strongest competing method on each benchmark.
- Caption memory helps historical QA (Finding 1). With the same checkpoint and Recent-16 window, retaining generated caption records raises OVOBench-Backward ASI from 63.5 to 71.6 and EPM from 62.0 to 62.6 relative to FIFO.
- PHCM improves perception and memory jointly (Finding 2). Using the same Recent-16 window as FIFO, PHCM raises OVOBench Real-Time from 80.9 to 81.4 and StreamingBench Real-Time from 86.3 to 86.9, while beating the Full-history variant on ASI by 4.0 points (71.6 vs. 67.6) and reaching the highest OVOBench Overall score of 72.1 (Full 70.9, FIFO 70.1). Its EPM (62.6) is slightly below Full (63.0).
- PSTL beats dense state supervision (Finding 3). Trained on the same data and sequences, PSTL supervises only 27.5% of annotated state tokens yet achieves 48.7 on ProactiveVQA, 36.6 on OmniMMI, and 41.6 on OVO-Timing, versus 47.0, 32.8, and 26.5 for dense supervision with frequency-balanced focal loss at 100.0% supervision, and versus 26.6, 27.4, and 1.4 for a supervision-matched random baseline at 27.5%. Transition Only at 18.1% supervision reaches 46.7, 27.4, and 15.3.
- Modest answer-stage overhead. On a 360 s OVOBench sample measured on a single NVIDIA H200, PHCM uses 9.98 GB GPU memory, 4,308 context tokens, and 0.124 s TTFT, versus 9.69 GB / 3,036 tokens / 0.094 s for FIFO and 25.18 GB / 62,094 tokens / 4.560 s for the full visual history. PHCM reduces context length by 93.1% and GPU memory by 60.4% relative to Full, and adds only 1,272 context tokens, 0.29 GB, and 0.030 s over FIFO.
- Qualitative behavior. In one demo, PHCM records details about a yellow table and books and later answers "The small yellow table upstairs." after the relevant frames have left the recent visual window; in another, the model stays in
</Silence>between target events and emits</Response>as new evidence supports an updated count, progressing from one to three.
Methodology in Plain English
OneStreamer processes a live video stream as temporally ordered clips. A vision encoder extracts visual tokens, a projector maps them into the LLM's embedding space, and a Recent-N FIFO sliding window keeps only tokens from the latest N observed frames. Those visual tokens are interleaved with a text history containing accumulated caption records and dialogue, so the sequence contains only information available at the current time.
The model predicts task-specific control tokens and then generates text when a predicted state initiates output. </Observe> starts a local-detail caption, </Summary> starts a semantic summary of a completed event, </Standby> signals that relevant evidence is emerging but insufficient, </Response> starts a user-visible answer, and </Silence> means continued observation with no caption or response. Generated records are appended to the text history in generation order and remain available after their source frames leave the window, so memory is textual rather than stored visual features — no external retriever, memory encoder, or access to discarded visual features is used.
Training targets are built so that each target is supported by evidence available at its assigned time. The caption branch releases local descriptions only after their supporting intervals have been observed and segment summaries only after the corresponding event is complete. The QA branch localizes a coarse evidence interval, verifies it by having a VLM answer from the clip alone, and then calibrates the response time by sliding windows, scoring reasoning-oriented QA by the conditional likelihood of the reference answer given each prefix and perception-oriented tasks by visual-text similarity, choosing the earliest timestamp that satisfies the reliability criterion.
For state supervision, PSTL groups state tokens by the transition between consecutive control states and sets a per-sequence quota equal to the largest group targeting an output anchor. Groups within the quota keep full supervision; larger groups are uniformly subsampled without replacement to match the quota. Unselected state tokens stay in the causal sequence and can condition later predictions, but they do not contribute to the state loss; caption and task text keep full next-token supervision. PSTL is not applied at inference — the model generates states and text from its learned distributions with no external transition policy.
OneStreamer is initialized from Qwen3-VL-4B-Instruct and trained in a single supervised fine-tuning stage on OneStreamer-1M plus offline data sampled from LLaVA-Video. The vision encoder is frozen; the multimodal projector and LLM are optimized for one epoch with a maximum sequence length of 131,072 tokens and a peak learning rate of 1e-5, on 32 NVIDIA H200 GPUs.
Why This Matters
The paper reframes proactive generation not just as producing replies but as a shared learning interface that connects perception, memory formation, and response timing inside one causal model. Instead of choosing between retaining history and keeping real-time perception sharp, it converts observations into compact, time-grounded text that remains usable after the pixels are gone — and it shows this can be done with selective state supervision rather than labeling every state token.
Real-world applications:
- Live-stream and broadcast assistants that must comment on events as they unfold and answer questions about moments that have already scrolled past.
- Wearable and egocentric agents that continuously record locally useful facts (objects, actions, scene context) for later questions.
- Security and monitoring systems, where an initially unremarkable event may become relevant later and alerts should fire when evidence actually arrives rather than at the end of a clip.
- Procedural and instructional assistance, where a model guides a user step by step and must decide when to speak based on what has been observed so far.
Industry relevance: the paper targets the practical constraints of deployed streaming models — limited context capacity, GPU memory, and time-to-first-token. Its reported reduction of context length by 93.1%, GPU memory by 60.4%, and TTFT from 4.560 s to 0.124 s relative to retaining full visual history, at gains over an 11B competitor, points to a route for serving long-lived video interactions on smaller models. The dataset and synthesis pipeline, and the source-video-level decontamination protocol, are also directly reusable for building future streaming training corpora.
Future Directions
- Beyond the two caption granularities. PHCM currently uses dense local-detail captions and event-level summaries; it is an open question whether additional levels or schemas would extend usable memory further.
- Generalizing the state-token design. PSTL is defined per task family with task-specific control-state vocabularies and output anchors; how these vocabularies should be unified or extended to new interactive tasks is not settled.
- Memory growth over very long streams. Caption memory is appended in generation order and never revisited; how record accumulation behaves over much longer horizons than the 360 s efficiency sample is not reported.
- Extending evaluation coverage. The reported results span eight benchmarks for perception, memory, and proactive response; the paper does not report results for all competing methods on all benchmarks (several table entries are left blank), so broader head-to-head coverage remains open.
Target Audience
Researchers and engineers working on streaming and online video-language models, long-context multimodal memory, and proactive or interactive assistants. It is also relevant to practitioners building deployable video agents who care about context budget, GPU memory, and latency, and to dataset builders interested in streaming caption and QA synthesis with evidence-aligned timing. Readers without background in video LLMs and autoregressive control-token decoding will find the method sections demanding.
Authors’ abstract
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.