Research
Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation Overview Research area: Robotics — vision-language-action (VLA) models, real-time robot control, Flow Matching ac

- arXiv
- 2609.39822
- Published
- 2026-09-30
- Authors
- Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, Lingfeng Zhang, Jianglin Zhang, Tao Zhang
AI summary
Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level EvaluationOverview
Research area: Robotics — vision-language-action (VLA) models, real-time robot control, Flow Matching action generation, and distributed inference/execution systems.
Technical level: Advanced (assumes familiarity with Flow Matching, diffusion-style denoising, action chunking, and robot control loops).
Scope: This technical report measures where end-to-end latency actually accrues in a VLA robot system (using π0.5 as the baseline), proposes a two-stage non-uniform denoising scheme that cuts action-generation steps from 10 to 2, and evaluates six real-time execution strategies plus the proposed sampler on a long-horizon bimanual garment-folding task.
What This Paper Is About
VLA models run slowly relative to the high-rate control that physical robots need, so actions generated from stale observations can keep driving a robot whose state has already changed. The paper first quantifies this timing gap by measuring both model-side inference cost and robot-side delays (camera, proprioception, command response), then attacks the model-side bottleneck with a two-stage denoising schedule that reuses standard Flow Matching for a long first step and SnapFlow-style shortcut refinement for a short terminal step, and finally builds a distributed framework to compare real-time execution methods under one common setup.
Key Contributions
-
A systematic latency-measurement procedure for both model and robot layers. Using π0.5, the authors analyze visual-language conditioning and Flow Matching generation, and measure camera timestamp offset, image readout, proprioceptive feedback, and control response to quantify end-to-end VLA timing.
-
A two-stage non-uniform Flow Matching denoising strategy. Stepwise analysis reveals stage heterogeneity in denoising; non-uniform integration combines standard Flow long-interval generation with SnapFlow short-interval terminal refinement, reducing NFE from 10 to 2 and mean model-inference time from 61.557 ms to 21.956 ms (a 2.804× speedup).
-
An open-source distributed real-time VLA inference and execution framework. Modular, thread-decoupled components jointly manage observation acquisition, policy inference, action buffering, temporal alignment, local control, and runtime logging, with independent inference, action-publication, and robot-control rates.
-
A unified comparative evaluation of six real-time execution methods (Naive Asynchronous Execution, Temporal Smoothing, Inference-time RTC, Training-time RTC, Legato, and VLASH) plus combinations of two-stage denoising with representative execution methods, on a shared robot platform, task, and runtime configuration.
Main Findings
-
Non-model latency is substantial. Measured on the physical robot: primary-camera timestamp offset 17.7 ± 1.8 ms (17.690 ms), primary-camera image readout 32.8 ± 3.3 ms (32.817 ms), proprioceptive feedback 26.4 ± 6.2 ms, and command-to-motion response 15.3 ± 5.4 ms. Standard Flow at 10 NFE measured 61.6 ± 6.2 ms across 620 inference calls.
-
The denoising velocity field is stage-heterogeneous. Across six velocity-field trajectories under standard 10-NFE sampling, RMS magnitude stayed broadly stable in early and intermediate stages with small angles between adjacent velocity fields, while near the endpoint most trajectories showed more pronounced magnitude changes and all six angle trajectories rose at the final step.
-
Two-step non-uniform denoising gives a large speedup. Inference time fell from 61.557 ms (10 NFE) to 21.956 ms (2 NFE) — a 2.804× speedup and 64.33% time reduction — on an NVIDIA GeForce RTX 4090 D (24 GB) in JAX, averaged over 620 inference calls after excluding the first JIT compilation.
-
Offline action error stays in the same order of magnitude. On two fixed validation sequences, joint MAE/RMSE and bimanual TCP error traded advantages between standard Flow and the two-stage sampler. Sequence I: Standard Flow joint MAE 0.008525 rad, RMSE 0.012877 rad, Left TCP 11.396 mm, Right TCP 7.676 mm; Two-stage 0.008396 rad, 0.013329 rad, 11.493 mm, 6.461 mm. Sequence II: Standard Flow 0.006503 rad, 0.008931 rad, 6.930 mm, 5.684 mm; Two-stage 0.006797 rad, 0.009104 rad, 4.831 mm, 6.718 mm.
-
Legato led overall task performance. Legato: 29/30 (96.7% success, 95% Wilson interval [83.3, 99.4]), 73.56 s mean completion time, 47.31 h⁻¹ throughput, at 2 Hz inference, 50 Hz publication, 5 NFE.
-
VLASH ranked second. VLASH: 28/30 (93.3%, [78.7, 98.2]), 77.73 s, 43.22 h⁻¹, at 2 Hz inference, 30 Hz publication, 10 NFE.
-
Temporal Smoothing was the best training-free method. 23/30 (76.7%, [59.1, 88.2]), 93.93 s, 29.38 h⁻¹, at 3 Hz inference, 30 Hz publication, 10 NFE.
-
Training-time RTC, Naive Asynchronous, and Inference-time RTC were weakest. Training-time RTC: 19/30 (63.3%, [45.5, 78.1]), 129.37 s, 17.62 h⁻¹ (1 Hz, 30 Hz, 10 NFE). Naive Asynchronous: 19/30 (63.3%, [45.5, 78.1]), 124.13 s, 18.37 h⁻¹ (1 Hz, 30 Hz, 10 NFE). Inference-time RTC: 18/30 (60.0%, [42.3, 75.4]), 133.73 s, 16.15 h⁻¹ (3 Hz, 30 Hz, 10 NFE). A higher replanning rate alone did not guarantee better task performance.
-
Handover continuity separates the methods. On joint left_j3, Legato had the lowest Tail Gap (0.0193 rad) and Switch Gap (0.0062 rad); VLASH followed at 0.0459 rad and 0.0173 rad. Naive Asynchronous reached maximum acceleration of 744.1 rad/s² with Tail Gap 0.1799 rad and Switch Gap 0.1538 rad. Temporal Smoothing had the lowest maximum acceleration (25.6 rad/s²) but larger gaps (0.0967 rad, 0.0998 rad). Inference-time RTC: 65.55 rad/s², 0.1310 rad, 0.1164 rad. Training-time RTC: 392.8 rad/s², 0.1201 rad, 0.0630 rad.
-
Local trajectory windows show distinct failure modes. Naive Asynchronous jumped near 107 s from −1.0 rad to −1.8 rad with velocity falling to approximately −25 rad/s, reaching 744.09 rad/s² and 0.8552 rad maximum absolute tail deviation; Temporal Smoothing reduced these to 10.27 rad/s² and 0.6751 rad. Inference-time RTC showed a discontinuity near 13.7 s (65.55 rad/s², 0.1813 rad); Training-time RTC showed a velocity reset near 84–87 s (197.16 rad/s², 0.8382 rad). VLASH: 34.58 rad/s², 0.4615 rad. Legato kept velocity mostly within [−1, 1] rad/s at 18.28 rad/s² and 0.4588 rad.
-
Smoothness metrics are not interchangeable. Maximum acceleration reflects local smoothness, while Tail Gap and Switch Gap characterize tail deviation and inter-chunk handover continuity; suppressing switching impulses and reducing inter-chunk trajectory deviation are related but distinct objectives.
Methodology in Plain English
The authors start by treating real-time VLA as an end-to-end timing problem rather than a pure model-speed problem. They build a calibration rig: a robot performs small-amplitude periodic sinusoidal joint motions while a display shows a host time code and phase bar. A RealSense camera records the time code and an end-effector ArUco marker, while the host logs camera timestamps, image receipt times, joint commands, and SDK feedback. Comparing timestamps and signal phases separates camera timestamp offset, image readout, proprioceptive feedback, and motion-response latency. Calibration uses visual acquisition at 30 frames/s and state recording at 100 Hz.
For the model side, they decompose inference time into a condition-encoding cost that does not scale with sampling steps, per-step Flow Matching costs that do, and a fixed post-processing cost. Then they inspect the velocity field produced at each of the 10 denoising steps — measuring its RMS magnitude and the angle between consecutive velocity fields at each evaluation — and find that most of the meaningful directional correction and magnitude change is concentrated near the endpoint, not in early or middle steps.
That observation motivates the proposed schedule: instead of splitting the 10 steps evenly, spend one network evaluation on a long interval (length 0.7) for coarse generation and one on a short terminal interval (length 0.3, τ = 0.3) for refinement. Training randomly picks, with equal probability at the batch level, either the standard Flow Matching branch or a two-stage branch. In the two-stage branch, the terminal step is supervised using a SnapFlow-style shortcut target built from two stop-gradient local velocity evaluations within the terminal interval, with the first-stage loss weight fixed at 1.0 and the terminal refinement weight λ_sc = 0.1. The student then predicts the mean velocity from τ directly to 0. At deployment, only two action-expert evaluations are performed, and no teacher branch is used.
On the system side, they separate observation acquisition and high-rate control on the robot from GPU policy computation. A Collector/RobotIO aggregates an overhead camera, two wrist cameras, and robot state into a continuously refreshed cache; an inference thread snapshots it and packages images, language, state, sampling steps, and latency fields into a request. Communication uses a TCP socket or WebSocket across hosts, or shared memory on the same host. Results return as action chunks with latency metadata, and the client computes each chunk's expected effective index before handling handover, blending, and indexing in the action buffer. A separate publication thread emits commands at the control rate.
They then instantiate six real-time execution strategies as pluggable modules sharing observation, communication, buffering, and control paths, and evaluate them on bimanual T-shirt folding with π0.5 as the base policy on two Agilex Piper manipulators (12 arm-joint dimensions plus two gripper dimensions, 640×480 RGB at 30 frames/s). Three T-shirts differ in size and color (small purple, large pink, medium red); each method runs 10 trials per garment, giving 30 trials per method and 180 physical trials total. Task success, mean completion time, successful-task throughput (R_succ × 3600/T̄), 95% Wilson intervals, Tail Gap, Switch Gap, and maximum acceleration are all reported. The experiments are not paired, so no between-method significance tests are conducted. Synchronous inference and Temporal Ensembling serve only as reference configurations for framework validation and are excluded from the unified physical comparison.
Why This Matters
Impact on research. The paper reframes real-time VLA as a joint model-plus-system timing problem, provides reusable latency-measurement instrumentation, and shows that where you spend denoising steps matters as much as how many you use. Its observation that an inference-time RTC method running at 3 Hz still underperformed methods running at 1–2 Hz is a direct challenge to the assumption that faster replanning automatically yields better task outcomes.
Real-world applications:
- Garment handling and textile automation — the evaluated task is bimanual T-shirt folding with wrinkled, variable-size, variable-color garments.
- Other long-horizon bimanual manipulation where inter-chunk continuity and handover smoothness dominate quality, such as packing, kitting, or assembly.
- Robotic platforms with slow perception chains — any system where camera readout (32.8 ± 3.3 ms) and proprioceptive feedback (26.4 ± 6.2 ms) meaningfully delay the observations a policy conditions on.
- Edge or on-premise robot deployments where shrinking inference from 61.557 ms to 21.956 ms at 2 NFE frees compute budget or enables higher control rates.
Industry relevance. The distributed framework's pluggable execution strategies, action-provenance logging, and independent inference/publication/control rates map directly onto production robot stacks that must debug handover discontinuities in the field. The open-source release (project page and GitHub repository) and the decision to hold base model, training data, robot platform, and runtime configuration constant address the field's comparability problem, where results are otherwise hard to compare across different base models and hardware.
Future Directions
- Closing the remaining gap. Combining two-step denoising with representative execution methods reduced inference cost with only a small reduction in task performance — quantifying and shrinking that residual performance drop is the obvious next step.
- Better allocation of the limited compute budget. The current implementation fixes τ = 0.3 and a 0.7/0.3 split; whether the split should adapt per observation, per task phase, or per action dimension is left open.
- Extending latency calibration. The paper notes that hardware-specific clock semantics, fixed offsets, and wrist-camera timing differences require separate calibration, so generalizing the measurement procedure across platforms and camera suites is unresolved.
- Beyond the two NFE settings tested. Whether the stage-heterogeneity principle extends to other sampling budgets, other action parameterizations, or other Flow-based action heads is not established here.
Target Audience
Researchers and engineers working on VLA models, robot policy inference, or real-time robot systems who need to understand where end-to-end latency actually comes from and how sampling-step allocation affects both speed and task success. It is most useful to readers already comfortable with Flow Matching, action chunking, and asynchronous inference/execution, and to practitioners who need a reference framework and measurement protocol for comparing real-time execution strategies on physical hardware. Beginners will find the plain-language framing of latency sources and the taxonomy of six execution methods accessible, but the denoising and training objectives require background knowledge.
Note on completeness: The provided content is truncated inside Section 5.2.2 ("Combination with Real-Time Execution Mechanisms"). The specific numerical results for combining two-stage 2-NFE denoising with Legato and Temporal Smoothing are therefore not reported in the available text; only the abstract-level statement that the combination "substantially reduces inference cost with a small reduction in task performance" appears.
Authors’ abstract
Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference cost, while robot-side delays mainly arise from perception acquisition, communication scheduling, and physical response. Analysis of the velocity field shows relatively stable magnitude and direction in early integration, followed by stronger directional correction near the terminal steps. Based on this stage heterogeneity, we propose two-stage non-uniform denoising, reducing the number of steps from 10 to 2 and model-inference time from 61.557 ms to 21.956 ms. We also develop a distributed real-time VLA framework with independent inference, action-publication, and robot-control rates, modular observation acquisition, and action-provenance logging. Using π0.5 as the baseline, we evaluate six real-time execution methods on a long-horizon physical garment-folding task. Legato performs best overall among training-based methods, while Temporal Smoothing leads among training-free methods; both perform strongly in task success, completion time, action continuity, and acceleration smoothness. Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance. These results motivate joint optimization of model-inference efficiency and robot-system timing.