Skip to content
AI.info

Research

RealtimeWAM: One-Step Asynchronous World Action Models

Overview Research area: Robot learning / computer vision — specifically World Action Models (WAMs), a family of robot manipulation policies that use video-generation backbones to guide action predicti

RealtimeWAM: One-Step Asynchronous World Action Models
arXiv
2610.06617
Published
2026-10-05
Authors
Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong, Shiqiao Gu, Shunzi Yang, Ruihao Gong, Shen Ren, Tianwei Zhang, Wenya Wang

AI summary

Overview

Research area: Robot learning / computer vision — specifically World Action Models (WAMs), a family of robot manipulation policies that use video-generation backbones to guide action prediction.

Technical level: Advanced. The paper assumes familiarity with diffusion/flow matching, denoising schedules, consistency distillation, Mixture-of-Transformers (MoT) architectures, and CUDA stream/event programming.

Scope: The paper introduces RealtimeWAM, a post-training framework that makes any MoT-based World Action Model generate actions in one denoising step and run its video and action experts concurrently, achieving roughly 14× to 25× end-to-end speedup on an NVIDIA H100 with under 1% success-rate loss across LIBERO, LIBERO-Plus, and RoboTwin 2.0.

What This Paper Is About

World Action Models predict future visual observations as an extra learning signal for robot control, but the efficient MoT variants still run slowly because the action expert performs many denoising steps and the video expert must finish before the action expert starts. The paper's goal is to remove both of those delays — making action generation one-step and allowing the two experts to execute at the same time — without meaningfully hurting task success rates.

Key Contributions

  1. RealtimeWAM, described as the first one-step World Action Model to achieve near-lossless performance relative to its multi-step counterparts, designed as a post-training variant applicable to any MoT-based WAM.
  2. Teacher-Anchored Consistency Distillation (TACD), which identifies a "local-global error gap" in standard consistency distillation — small local consistency error does not guarantee an accurate final action — and adds explicit supervision from the frozen teacher's multi-step rollout endpoint.
  3. Cross-Expert Wavefront Pipelining (CEWP), a block-level execution schedule that overlaps the Video and Action Experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it.
  4. Extensive evaluation across LIBERO, LIBERO-Plus, and RoboTwin 2.0, on two base architectures (Fast-WAM and Faster-WAM), showing under 1% accuracy drop and roughly 25× speedup on an H100.

Main Findings

  • Intra-expert iteration dominates latency: The action expert accounts for nearly 90% of end-to-end latency for action generation across diverse GPUs, and existing MoT-based WAMs such as Fast-WAM require over 300 ms for a single inference call on an H100.
  • RoboTwin 2.0 results: RealtimeWAM* (built on Fast-WAM) reaches 90.84% overall success with one-step action generation, a 0.67% reduction versus 10-step Fast-WAM. RealtimeWAM† (built on Faster-WAM) reaches 92.64%, only 0.29% below 10-step Faster-WAM. Both exceed other efficient WAMs (81.41% for one-step Flash-WAM) and step-distillation baselines (88.22% for DMD*).
  • LIBERO results: RealtimeWAM* achieves 97.0% overall with one-step generation, matching 10-step Fast-WAM and improving 2.0% over directly reducing Fast-WAM to a single denoising step (95.0%). It outperforms one-step Flash-WAM (95.1%) and DMD* (96.1%), and is competitive with Light-WAM (97.2%). RealtimeWAM† reaches 99.0%.
  • LIBERO-Plus robustness: RealtimeWAM† achieves 73.0% overall across the seven perturbation subsets, only 0.6% below Faster-WAM (73.6%), and slightly improves on camera, sensor-noise, and layout perturbations — steps that preserve out-of-distribution robustness.
  • Ablation on loss terms: With the Video Expert frozen, replacing consistency-only supervision (89.85% overall) with consistency plus the teacher-anchored loss (TACD) raises RoboTwin 2.0 success to 90.84%. Fine-tuning the Video Expert during distillation lowers the overall success rate and adds training overhead.
  • Teacher rollout steps: K = 5 gives 90.39% overall, K = 10 gives 90.84%, and K = 20 gives 90.77% while requiring more teacher computation, so K = 10 is the default.
  • Latency breakdown on H100: TACD alone cuts latency from 299.7 ms to 65.8 ms on Fast-WAM and from 218.9 ms to 57.1 ms on Faster-WAM (4.56× and 3.83×). After CUDA Graph, CEWP further reduces latency from 23.3 ms to 17.4 ms on Fast-WAM and from 26.2 ms to 22.4 ms on Faster-WAM (1.34× and 1.17×). With all optimizations, RealtimeWAM reaches 12.2 ms and 16.1 ms, for overall speedups of 24.55× and 13.56×.
  • Training overhead of the teacher anchor: Adding the teacher-anchored loss increases training time from 5h 21min to 8h 3min (about 50%), comparable to jointly fine-tuning both experts (8h 1min). Freezing the Video Expert cuts peak GPU memory from 65.98 GiB to 24.73 GiB versus joint fine-tuning, a 62.5% reduction.
  • Theoretical guarantee: The paper decomposes total endpoint error into local consistency error plus global target error, and shows the expected squared endpoint error is bounded by twice the teacher-anchored loss plus twice the expected squared teacher numerical integration error, for 0 < t ≤ 1.
  • Not applicable to shared-backbone WAMs: The framework targets MoT-based WAMs and is not directly applicable to shared-backbone architectures such as DreamZero; it does support both MoT models that omit explicit future-video prediction at inference (Fast-WAM) and those that retain it (Faster-WAM).

Methodology in Plain English

The starting point is an efficient WAM built from two transformer "experts": a Video Expert that reads the observation and instruction and produces a cached set of keys and values (the KV cache), and an Action Expert that uses that cache to produce a chunk of robot actions by repeatedly denoising random noise.

To cut the number of denoising steps, the authors use consistency distillation — training a student to jump straight from noisy actions to clean actions. They observe that this only enforces local agreement between the student and a moving-average target; it does not guarantee the student lands on the right final action. So they add a second loss term: the frozen teacher model runs a full multi-step rollout (10 steps) to produce an endpoint action, and the student's velocity prediction is directly matched against the teacher's interval-averaged velocity toward that endpoint. They set the weight of this term to 0.2, and train with LoRA (rank 128) on the Action Expert only, freeing the Video Expert. Training runs 30,000 iterations with AdamW, a learning rate of 10⁻⁴, weight decay 10⁻², and EMA decay 0.995.

To remove the waiting between experts, they analyze the actual data dependencies inside each transformer block. An action block needs only the video keys and values from its own corresponding block, not the Video Expert's complete output, and the action projection can run before that cache arrives. So they place the two experts on separate non-blocking CUDA streams and have the video stream record a CUDA event right after each block's KV projection; the action stream waits on that event only immediately before its attention step. This turns sequential expert execution into a staggered block-by-block "wavefront." They also employ CUDA Graph capture and efficient kernels from LightX2V.

Why This Matters

The work reframes WAM inference latency as two separable bottlenecks — repeated action denoising and coarse expert synchronization — and shows both can be attacked with post-training and scheduling changes rather than new architectures. Because RealtimeWAM is applied on top of existing models (Fast-WAM and Faster-WAM) without retraining from scratch, it offers a practical recipe for making existing WAM checkpoints deployable at control rates. The paper reports 12.2 ms per-call latency, which is shorter than typical execution times such as a 30 Hz control frequency.

Real-world applications implied by the evaluation:

  • Robot manipulation arms performing tabletop tasks such as the LIBERO suites (Spatial, Object, Goal, Long) and the 50 tasks of RoboTwin 2.0.
  • Dual-arm robotic systems, which the authors evaluate on real-world dual-arm manipulation tasks (details deferred to an appendix).
  • Real-time closed-loop control, where per-call latency below the controller's execution period avoids stalls and stale observations.
  • Deployment on lower-cost GPUs, since latency is also benchmarked on an RTX 4090D and RTX 5090.

Industry relevance: robotics and automation, particularly for companies shipping learned manipulation policies where inference cost and control-rate latency determine feasibility. The author affiliations include Sensetime and Continental Automotive Singapore, alongside Nanyang Technological University and Beihang University.

Future Directions

  • Extending beyond MoT architectures. The authors explicitly state the method is not directly applicable to shared-backbone WAMs such as DreamZero, where action-specific characteristics are harder to isolate from video step distillation.
  • Reducing teacher-rollout overhead. The teacher-anchored loss adds roughly 50% training time; cheaper or better-approximated teacher endpoints are an open avenue.
  • Broadening the distillation comparison. The paper notes that prior ultra-few-step distillation for WAMs still shows noticeable degradation on RoboTwin and real-world tasks, leaving room to test whether TACD closes that gap in settings that retain future-video denoising, such as LingBot-VA.
  • Wider hardware and real-world validation. Latency results for RTX 4090D and RTX 5090 and real-world dual-arm results are placed in appendices, leaving larger-scale real-robot evaluation as a natural next step.

Target Audience

Researchers and engineers working on robot learning, Vision-Language-Action models, and diffusion/flow-based generative policies, especially those concerned with inference latency and real-time deployment. It will also interest practitioners of few-step distillation and of systems-level optimization for multi-expert transformer architectures. Readers without a background in diffusion sampling or transformer internals will find the theory and scheduling sections demanding, since the paper is written at an advanced research level.

Authors’ abstract

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, $&lt;1\%$ drop) across these benchmarks while delivering significant end-to-end speedup (\eg, $\sim25\times$ on H100). Our code and checkpoints are available via this \href{https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam}{link}.

Read the original paper