Research
Rolling-WAM: World Action Models with Rolling Imagination
Overview Research area: Robotics — robot manipulation policies, specifically World Action Models (WAMs) that jointly predict actions and future video, with a focus on closed-loop inference efficiency.

- arXiv
- 2609.30247
- Published
- 2026-09-24
- Authors
- Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
AI summary
Overview
Research area: Robotics — robot manipulation policies, specifically World Action Models (WAMs) that jointly predict actions and future video, with a focus on closed-loop inference efficiency.
Technical level: Advanced. The paper assumes familiarity with diffusion/flow-matching sampling, Transformer architectures, and vision-language-action policy evaluation.
Scope: The paper proposes Rolling-WAM, a formulation that spreads the joint video-action denoising process across successive replanning cycles, and evaluates it on LIBERO, RoboTwin 2.0, and three real-world Unitree G1 humanoid tasks.
What This Paper Is About
World Action Models predict robot actions together with future video, which gives the policy visual context about how a scene may evolve. The problem is that standard joint samplers denoise the entire prediction horizon from pure noise at every replanning cycle, so inference is slow and action updates arrive late. Rolling-WAM reorganizes this computation: instead of denoising a whole new horizon each time, it keeps a sliding window of video-action chunks at staggered noise levels and fully denoises only the imminent chunk, while partially refining farther-future chunks that are retained and refined again in later cycles.
Key Contributions
-
A rolling formulation for joint video-action models. Rolling-WAM distributes joint denoising across successive replanning cycles, maintaining a sliding window of aligned video-action chunks at progressively higher noise levels toward the future.
-
Two operational modes. A rolling mode schedule, σ_j^rolling(τ) = σ((j−1+τ)/W), guarantees that after the first chunk is removed, every retained chunk already sits at the starting noise level for its new position; an initialization mode schedule, σ_j^init(τ) = σ(min{1, τ + (j−1)/W}), builds the window from Gaussian noise at episode start.
-
An architecture that lets actions attend to partially denoised futures. A pretrained video Diffusion Transformer is paired with a lightweight action Transformer in a Mixture-of-Transformers design, with masked joint attention that lets every action chunk attend to visual features across the entire prediction window (action-to-action attention stays within a chunk, and video tokens do not attend to action tokens).
-
A speedup with competitive performance. The method delivers a 4.5× steady-state replanning speedup over standard joint WAMs, while reaching 98.1% average success on LIBERO, 93.3% on RoboTwin 2.0, and 85.0% average on real-world humanoid tasks.
Main Findings
-
LIBERO: Rolling-WAM reaches 98.1% average success (98.2 Spatial, 98.0 Object, 98.2 Goal, 97.8 Long), within 0.4 percentage points of the joint leaders LingBot-VA and Joint-WAM (both 98.5%), and above Motus (97.7%) and Fast-WAM (97.6%). Success ranges from 97.8% to 98.2% across the four suites.
-
RoboTwin 2.0: Rolling-WAM achieves 93.5% in the Clean setting and 93.0% in the Randomized setting, for a 93.3% average that leads LingBot-VA (92.2%), Fast-WAM (91.8%), and Joint-WAM (90.6%) — without embodied pretraining. Per-task results on all 50 tasks are given in the paper's Table V.
-
Real-world humanoid (Unitree G1): Rolling-WAM attains the highest average success rate of 85.0% among five policies, ahead of Joint-WAM (78.3%) and Fast-WAM (75.0%). It matches the strongest baselines on Doll Placement (85%) and Plate Stacking (100%), and achieves the highest rate on Bead Pouring (70%).
-
Latency at default settings: With N = 10 denoising steps per chunk and W = 5 chunks, Rolling-WAM uses two denoising steps per update and takes 215 ms, versus 978 ms for Joint-WAM and 548 ms for Fast-WAM — roughly 4.5× and 2.5× speedups. The VLA baselines π0.5 and GR00T N1.7 take 296 ms and 285 ms.
-
Longer horizon at lower latency: Rolling-WAM retains an 80-action prediction window while the baselines predict 16 actions, so the latency reduction comes while refining a longer future at each update.
-
Window-size scaling: When W is varied from 1 to 8, at W = 5 Rolling-WAM requires 322 ms versus 832 ms for Fast-WAM and 1529 ms for Joint-WAM (2.58× and 4.75× speedups). The latency curve flattens for W = 6–8, where the number of rolling denoising steps stays at two per replanning cycle.
-
Window-size ablation on task success: On six selected RoboTwin Clean tasks, success peaks at 78.2% for W = 5, versus 77.3% for W = 3 and 69.5% for W = 8. Larger windows do not consistently help.
-
Training noise mixture ablation: The full rolling schedule gives 78.2% on the selected RoboTwin tasks and 98.1% on LIBERO. Random sampling with p = 0.5 raises RoboTwin to 78.5% but lowers LIBERO to 97.3%; no alternative improves both benchmarks.
-
Cross-chunk action attention ablation: Allowing direct action-to-action attention across chunks gives 76.3% on selected RoboTwin tasks and 97.9% on LIBERO, compared with 78.2% and 98.1% when action attention is restricted within each chunk.
-
Qualitative behavior: Comparing imagined and observed rollouts on Put Object Cabinet, imagined robot motion and task progression closely follow observations. On the humanoid, execution appears more continuous across action-chunk boundaries than with Joint-WAM, where pauses are more apparent.
Methodology in Plain English
The policy takes a camera observation, robot state, and a language instruction, and must output a sequence of actions. Rolling-WAM splits the future into W chunks, each containing K actions plus the matching video segment, for a total horizon H = W·K. Each chunk carries its own noise level, increasing toward the far future.
At each replanning cycle, the model runs only N/W denoising steps, which fully clean the nearest chunk so it can be executed. After the robot executes that chunk, the window slides: the executed chunk is dropped, the remaining partially denoised chunks are retained, and a fresh chunk initialized with Gaussian noise is appended. Because the noise schedules are constructed so that a chunk sliding one position forward lands exactly on its required starting noise, retained predictions continue refining rather than being recomputed. A new camera observation and state condition the next cycle.
The model itself uses flow matching. A pretrained video Diffusion Transformer (initialized from Wan2.2-TI2V-5B, with its text encoder and VAE reused) and a 30-layer, 1024-dimensional action expert (approximately 1B parameters, initialized by interpolating the video expert's weights) are coupled by masked joint attention. Language and proprioceptive embeddings enter through cross-attention. Video and action tokens are modulated by their respective chunk-wise noise levels so different refinement stages can be processed together. Training samples τ uniformly and picks rolling or initialization mode with probabilities 0.8 and 0.2 respectively, using a shifted noise schedule with ρ = 5.
Evaluation covers the Spatial, Object, Goal, and Long suites of LIBERO (10 tasks and 500 demonstrations each, one policy for 10 epochs, effective batch size 128, 50 rollouts per task); 50 bimanual tasks in RoboTwin 2.0's Clean and Randomized settings (2,500 clean and 25,000 randomized demonstrations, 5 epochs, effective batch size 1024, 100 rollouts per setting); and three Unitree G1 tasks — Doll Placement, Plate Stacking, and Bead Pouring — using a 320×224 egocentric RGB image and a 43-dimensional state, predicting 78-dimensional actions (a 64-dimensional SONIC motion latent plus 7-dimensional commands per hand), trained with 50 demonstrations per task at 10 Hz for 7,500 steps at effective batch size 192, and evaluated with 20 trials per task. Latency was measured on a single NVIDIA A100 under the RoboTwin 2.0 setting at 384×320 resolution, with timing including visual encoding and denoising under CUDA synchronization and excluding warm-up and initialization, and with no torch.compile, TensorRT, or custom CUDA kernels.
Why This Matters
Impact on research: The paper reframes denoising cost as something to be scheduled across control cycles rather than paid in full at every replan. It extends rolling diffusion from action-only buffers to coupled video-action generation, so action tokens can condition on partially denoised visual futures. The authors note the approach is orthogonal to other WAM efficiency techniques, such as caching or token reduction, and can be combined with them.
Real-world applications:
- Humanoid manipulation in homes and warehouses, where a robot must respond quickly to changing scenes while still planning ahead.
- Bimanual tabletop assembly and rearrangement, as measured in the RoboTwin 2.0 tasks.
- Any deployment running on modest hardware, where the 4.5× speedup reduces the compute budget needed for closed-loop control.
- Long-horizon tasks where retaining a longer horizon (80 actions predicted per update versus 16 for the baselines) improves consistency across chunk boundaries.
Industry relevance: Latency is a practical barrier to deploying generative world models on physical robots. A method that cuts per-update latency below that of the compared VLA baselines while still generating future video — and doing so without torch.compile, TensorRT, or custom kernels — is relevant to teams building real-time robot policies.
Future Directions
-
Window and chunk size selection. The authors state that design choices such as window and chunk sizes warrant further exploration across task settings with different dynamics and control frequencies.
-
Handling rapid scene change. The paper notes that, despite observation feedback, retained predictions may lag behind rapid scene changes and misguide action generation, particularly with long windows.
-
Adaptive window management and asynchronous execution. Both are offered as directions for improving closed-loop responsiveness.
-
Combining with other efficiency techniques. The paper explicitly frames rolling denoising as orthogonal to complementary model and system optimizations for low-latency robot control.
Target Audience
Robotics and embodied-AI researchers working on world models, diffusion/flow-matching policies, and vision-language-action models; engineers deploying learned manipulation policies on real hardware under latency constraints; and graduate students or advanced practitioners who already understand diffusion sampling and transformer architectures and want to see how inference can be restructured for closed-loop control.
Authors’ abstract
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.