Skip to content
AI.info

Research

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models Overview Research area: Computer Vision / robot learning — specifically memory for vision-language-action (VL

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
arXiv
2609.05533
Published
2026-09-02
Authors
Cheng Yin, Wang Xu, Junpeng Yang, Sikyuen Tam, Hanyu Liu, Yuan Yao, Xiangrui Zeng, Junbo Cui, Yequan Wang, Zhouping Yin, Yankai Lin

AI summary

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Overview

Research area: Computer Vision / robot learning — specifically memory for vision-language-action (VLA) models operating in long-horizon, partially observable manipulation tasks.

Technical level: Advanced. The paper assumes familiarity with VLA policies, vision-language model (VLM) backbones, flow-matching action heads, key-value caching, and attention mechanisms such as sliding-window attention.

Scope: The paper proposes and evaluates SimpleMemVLA, a VLA architecture that removes the dedicated memory module entirely and instead feeds a timestamped window of past camera frames directly into a pretrained VLM backbone's native video channel, then tests this design across six manipulation suites.

What This Paper Is About

Long-horizon robot manipulation is partially observable: the observation that determines the correct next action may have appeared minutes earlier, so two moments with identical current camera images can require different actions. Existing VLAs handle this with dedicated memory machinery — retrieval banks, learned compressors, or recurrent states — but all of these must decide what to keep from the past before knowing what a future decision will need, a problem the authors call "write-time commitment." The paper's goal is to show that at minute-scale histories, no such machinery is needed: the history fits inside the backbone's context window and can be read directly by native attention.

Key Contributions

  1. A VLA with no dedicated memory module. SimpleMemVLA keeps the sampled history intact and passes it to the backbone (Qwen3.5-4B) as a plaintext-timestamped video clip, letting the model's native video attention select relevant past evidence at read time rather than committing at write time.

  2. A narrow text channel from history to action. The generated textual sub-task span is the only channel carrying historical information into a standard flow-matching action head: the action expert conditions on the contextual hidden states fused with the token embeddings of that span, plus a single-token encoding of the normalized proprioceptive state. The expert receives no prompt tokens directly.

  3. Exact streaming inference for deployment. Because consecutive decisions share almost their entire video-history prefix, the shared prefix is prefilled during action execution and its key-value cache reused, reducing decision latency from 1.02 s to 0.68 s with identical outputs to full recomputation.

  4. A matched-stack attribution study. Retrieval, token-compression and recurrent-state mechanisms are re-implemented on the otherwise unchanged SimpleMemVLA stack (same training data, backbone, sub-task supervision, action head and optimizer) to isolate the memory interface as the cause of the gains rather than the backbone.

Main Findings

  • State of the art on all four memory benchmarks. SimpleMemVLA achieves the highest reported performance on RMBench, RoboMME, MIKASA-Robo and RoboMemArena, while tying the best reported average on LIBERO and achieving the strongest reported zero-shot transfer to LIBERO-Plus.

  • RMBench: 94.0% with a single multi-task model. On the nine RMBench tasks with published baselines, this is 11.0 points above the strongest specialist, even though every baseline trains a separate task-specific specialist. It is the only method in that table whose average rises from the single-memory setting to the multi-memory setting, from 91.6% to 97.0%. Place-Mat has so far been evaluated only by SimpleMemVLA, so it is excluded from all cross-method averages.

  • RoboMME: 88.3%. This is 43.7 points above the strongest non-oracle baseline, first on all sixteen tasks, and above the GroundSG reference supplied with ground-truth perception (84.1%). The authors trace that gap to write time, noting that even with ground-truth perception performance reaches only 84.1%.

  • Native context beats every re-implemented memory family on the matched stack. On RoboMME, native context reaches 88.3%, while retrieval reaches 31.5%, token compression 22.6%, and recurrent state 20.6%. At the task level, retrieval is the strongest alternative because it retains raw visual content and reaches 42.0% on Permanence, but presenting frames without order or timestamps reduces the swap and ordered-pattern tasks to 10–20%. Token compression observes every history frame yet tops out at 46.0% on Counting; the recurrent variant tops out at 31.0% on Counting and falls to 24.0% on Permanence, 17.5% on Reference and 10.0% on Imitation.

  • How much of the original history survives matters. Retrieval exposes eight uniformly sampled frames as unordered images without timestamps; token compression maps the full history into a fixed 64-token representation using queries independent of the current decision; recurrent state continually rewrites the past into a fixed 16-token representation. Normalized to native-context performance, the three variants retain at most 79%, 96% and 58% respectively, with the token-compression peak confined to the count task SwingXtimes.

  • MIKASA-Robo: 74.0%. This is 29.6 points above the strongest prior VLA and 6.2 points above GMP (67.8%), a memory-specialized non-VLA policy trained separately per task. Because the benchmark shows and removes its cues before the robot may act, all evidence exists only in history.

  • RoboMemArena: 63.6% TSR and 72.1% CSR with one model and a 126 s native context window, which is +17.4 points over the strongest prior model and above the benchmark's own ground-truth reference (46.1%). The gain concentrates on memory-dependent categories — Occlusion rises from 39.1% to 64.3% TSR and Counting from 31.4% to 71.4% compared with FrameSamp+Modul — and the method does not lead on Transferring.

  • No cost on general-purpose control. SimpleMemVLA reaches 97.5% across the four LIBERO suites (500 trials per suite), tying the best reported average, and transfers zero-shot to the 10,030 perturbed LIBERO-Plus tasks at 78.4%, 5.3 points above the strongest memory-augmented model. RIPT-VLA ties it on standard LIBERO yet drops 29.1 points under perturbation, against 19.1 for SimpleMemVLA.

  • The policy genuinely reads its history. Masking task-relevant evidence changes the policy's outputs, whereas masking an irrelevant segment does not. The policy also adapts to edited or previously unseen visual histories without parameter updates, which the authors describe as a form of visual in-context learning.

  • Context is cheap at this scale. A 60 s history occupies roughly 5.6k tokens of the backbone's 262k-token context window, which can hold roughly 45 minutes.

Methodology in Plain English

The design rests on a simple observation: modern VLM backbones were pretrained on timestamped video, and minute-scale robot history is small enough to feed them directly. So instead of building a memory system, SimpleMemVLA rebuilds a sliding window of head-camera frames at every prediction step, subsampled well below the native frame rate, and hands it to the backbone through its regular video channel — with each temporal patch labeled by its offset within the window. Current wrist-camera views enter through the image channel without timestamps, following one rule: multi-frame cameras become video, single-frame cameras become images, so the input format itself separates past from present.

The backbone then writes a one-sentence description of the current sub-task, decoded as an ordinary assistant response under an unmodified chat template and capped at 64 tokens. That answer span, not the raw history, is what the action expert sees: its hidden states fused with its token embeddings condition a DiT-style flow-matching head that generates an action chunk from Gaussian noise via a few Euler steps. Training combines token-level cross-entropy on the sub-task target — generated offline by a cloud LLM from each demonstration — with the flow-matching objective on demonstrated action chunks. The same sub-task targets supervise all three mechanism variants, so the comparison isolates the memory interface rather than annotation quality.

For deployment, the authors exploit the fact that consecutive decisions differ by only one temporal patch: they prefill the shared prefix while the robot executes the current chunk and reuse the cached keys and values, so only the newly arrived patch and the instruction are processed at the next decision. Embodiment-specific choices live entirely in a configuration tuple specifying camera sets, window length, sampling rate, action horizon and action dimensionality, so moving between bimanual and single-arm platforms changes configuration rather than the memory mechanism.

Why This Matters

Impact on research. The paper reframes memory in VLA as a read-time selection problem rather than a write-time compression problem, and supports that reframing with a controlled comparison in which the backbone, data, supervision and action head are held fixed. It also reports that a ground-truth-perception symbolic reference (84.1% on RoboMME) is outperformed by native context (88.3%), which locates the limitation of prior mechanisms in when they commit information rather than how accurately they extract it. The result echoes an analogous finding in streaming video understanding, where an off-the-shelf VLM given a sliding window of recent frames matches or outperforms dedicated streaming-memory methods.

Real-world applications:

  • Long-horizon household and tabletop manipulation where an object is revealed once and must be remembered after occlusion.
  • Multi-step assembly or electronics manufacturing tasks that require counting completed actions and tracking ordered sequences over many control steps.
  • Industrial bimanual platforms where the same model must serve several tasks without training a separate specialist per task.
  • Robust deployment under visual perturbation — camera changes, altered lighting, changed backgrounds — where a retained history acts as a stable visual reference.

Industry relevance. Training one multi-task checkpoint instead of one specialist per task reduces the engineering and compute burden of deployment, and the 0.68 s decision latency brings a minute-scale-memory policy close to the control-loop cost of a single-frame VLA. The code is released at https://github.com/wadeKeith/SimpleMemVLA.

Future Directions

  • Continual inference over unbounded streams. The current scheme handles bounded, minute-scale context; the authors explore a sliding-window-attention variant trained with episode-absolute timestamps to keep the active context and cache bounded as history grows.
  • When the history no longer fits. The paper explicitly targets minute-scale histories that fit inside the context window; the regime where full history exceeds the real-time budget is delegated to designs such as MEM, leaving open how native-context memory should degrade beyond roughly 45 minutes of context.
  • Training signal for state-update decisions. The authors note that in their recurrent implementation, training backpropagates through only four decisions, so supervision from later decisions cannot directly teach earlier state updates which observations should have been preserved — leaving the question of whether better credit assignment could rescue recurrent approaches.
  • Cross-embodiment generality. Since embodiment-specific choices are confined to a configuration tuple, the natural test is how far this configuration-only transfer extends across bimanual and single-arm platforms not covered in the six evaluated suites.

Target Audience

Robotics and embodied-AI researchers working on VLA policies and long-horizon manipulation; engineers deploying generalist robot policies under real-time constraints who need to weigh memory designs against latency; and multimodal-LLM researchers interested in how far native long-context video attention can substitute for purpose-built memory modules. Readers without background in VLA architectures, flow matching or transformer inference will find the experimental comparisons accessible but the method sections demanding.

Authors’ abstract

Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms: retrieval banks, learned compressors, recurrent states must decide what to keep from the past before knowing what a future decision will require. This was motivated by the assumption that minute-scale history is too large to process directly, which modern VLM backbones no longer make true. In this work, we introduce SimpleMemVLA, a VLA without a dedicated memory module. It keeps the sampled history intact and passes it to the backbone in the timestamped video format the backbone was pretrained to process; the hidden states of a generated sub-task then form the only channel from history to a standard flow-matching action head. Since consecutive decisions share most of their history, prefilling the shared prefix during action execution keeps latency close to a single-frame VLA. SimpleMemVLA sets a new state of the art on four memory benchmarks without cost on general-purpose control. Holding the backbone and training setup fixed, it outperforms retrieval, compression and recurrent-state mechanisms by a wide margin, and causal interventions confirm that the policy genuinely reads its history. Code available at https://github.com/wadeKeith/SimpleMemVLA

Read the original paper