Research
Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models
Overview Research area: Robotics — generative action policies for robotic foundation models (RFMs), specifically one-step (single forward pass) action generation to replace multi-step flow matching. T

- arXiv
- 2610.00864
- Published
- 2026-10-01
- Authors
- Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
AI summary
Overview
- Research area: Robotics — generative action policies for robotic foundation models (RFMs), specifically one-step (single forward pass) action generation to replace multi-step flow matching.
- Technical level: Advanced. The paper builds on flow matching, MeanFlow, Jacobian-vector products, and stop-gradient bootstrapping, and assumes familiarity with diffusion/flow-based policies.
- Scope: The paper diagnoses why MeanFlow collapses when applied to generalist RFMs, proposes a kinematic decomposition ("Kinematic MeanFlow", K-MF) that enables one-step action generation, and validates it across two RFMs, four datasets, two training paradigms, and both Sim2Sim and Real2Sim evaluation.
What This Paper Is About
Robotic foundation models typically generate action chunks with flow matching, which requires several iterative passes through a transformer action head — and that action head consumes a large share of total inference time (about 40% of end-to-end latency on an NVIDIA L40, and up to 50% on a Jetson Orin, per the paper's motivation section). MeanFlow is a promising recipe for collapsing this to a single step in image generation, but the authors find that applying it directly to generalist RFMs causes a performance collapse. The paper's goal is to explain that collapse and fix it, so that RFMs can produce actions in one step without losing accuracy.
Key Contributions
-
Diagnosis of MeanFlow failure in RFMs. The authors identify two distinctive dynamics in the RFM velocity field: local acceleration is stable early but surges sharply in the late denoising stage (below timestep t = 0.3, reaching 76× its value at t = 1), and the spread of its magnitudes across samples widens as denoising proceeds. These two effects corrupt MeanFlow's bootstrapping loop and produce an unstable, spiky loss landscape.
-
Kinematic MeanFlow (K-MF). Using a kinematic identity, K-MF decomposes the time-derivative term of the MeanFlow formulation into two sub-interval terms separated by an intermediate point c, weighted by (1−λ)² and λ². One term handles early-stage denoising dynamics, the other handles late-stage corrections, and the extra estimation at c acts as an anchor that mitigates error amplification.
-
Three strategies for choosing the intermediate point. Random sampling (λ ~ U(0,1)), fixed at the midpoint (λ = 0.5), and a learnable strategy where a lightweight 3-layer MLP predicts λ = G_φ(t, r), trained with an auxiliary objective that uses a weighting term ω with γ annealed linearly from −0.5 to 0.5.
-
Broad empirical validation plus efficiency analysis. Evaluation covers fine-tuning of GR00T-N1.6 and training from scratch with SimVLA-S, four datasets spanning 658 to 87K trajectories, and latency measurements on L40 and Jetson Orin in eager and compiled modes.
Main Findings
-
MeanFlow needs a delicate recipe and still fails at one step. On GR00T-N1.6 with LIBERO-10, MeanFlow with gradient clipping and progressive timestep sampling requires sufficient iterations to become viable: it scored 52.5% at 40K iterations, 87.0% at 60K, and dropped to 84.5% at 80K, while at one step (NFE=1) it failed outright. With gradient clipping and progressive sampling both off, it also failed at NFE=2.
-
Existing variants fall short at one step. MVP reached 82.0% at two steps but produced no viable actions at one step, and α-Flow reached 94.0% at two steps but dropped to 84.5% at one step. K-MF reached 94.5% at one step with 20K iterations, matching the 94.5% four-step flow matching baseline.
-
Consistent gains across training paradigms. On LIBERO, fine-tuned GR00T-N1.6 (3B) averaged 97.9% with K-MF at NFE=1 versus 97.3% for four-step flow matching and 94.5% for two-step MeanFlow. Trained from scratch, SimVLA-S (0.6B) averaged 95.1% with K-MF at NFE=1 versus 94.2% for ten-step flow matching and 93.7% for two-step MeanFlow.
-
Generalization across benchmarks and embodiments. On Fractal (Real2Sim, GoogleX), K-MF averaged 78.4% versus 73.9% for four-step flow matching; on BridgeData V2 (Real2Sim, WidowX) it averaged 59.9% versus 58.4%; on COMPASS-generated PointNav (Sim2Sim, Unitree G1) it averaged 74.0% versus 73.8%.
-
The learnable strategy is best; partition diversity is not the reason. The learnable strategy scored best on both models (97.9% on GR00T-N1.6, 95.1% on SimVLA-S), the fixed midpoint was second-best (97.6% and 94.4%), and random sampling was clearly weaker (95.3% and 93.8%). The learned λ correlates strongly with interval width t − r, suggesting the benefit comes from better time-derivative estimation rather than augmentation from different partitions.
-
Substantial inference savings. K-MF reduces action-head latency of GR00T-N1.6 by 67.5%–74.4% and end-to-end latency by 30.3%–54.9% across L40 and Jetson Orin in eager and compiled modes. In the flow matching baseline, the action head accounted for 43%–73% of end-to-end latency.
-
Modest training overhead. Relative to MeanFlow, K-MF adds 7.5%–13.6% training time and at most 2.5% peak GPU memory. Relative to flow matching, the training overhead is 33% for GR00T-N1.6 and 27% for SimVLA.
-
Velocity-field dynamics depend on the data. GR00T-N1.6 and SimVLA-S show a similar increasing trend in local acceleration on the same dataset; a model trained on the larger, more complex Fractal dataset shows a wider spread than one trained on LIBERO; and the relative spread within Fractal grows with task count, following an approximately logarithmic trend at low task counts and a saturating power-law trend at higher task counts.
Methodology in Plain English
The authors start from the observation that flow matching predicts an instantaneous velocity at each denoising timestep and therefore needs many passes, whereas MeanFlow predicts the average velocity between any two timesteps and can therefore jump from noise to action in one pass. MeanFlow trains by a bootstrapping loop: the model's own estimate of how the average velocity changes over time is used (with a stop-gradient) as part of the regression target.
The authors measured how quickly the average velocity changes (local acceleration) over the denoising process, using an interval of Δt = 0.05, and compared image generation (DiT-B on ImageNet) with action generation (SimVLA-S paired with DiT-B on LIBERO). In images this quantity stays flat with a narrow spread. In RFM action generation it stays stable early and then spikes late, and its spread across samples grows. That means the derivative the model must estimate becomes both large and highly variable near the end, so the bootstrapping loop feeds back badly estimated targets and training destabilizes.
Their fix is a mathematical identity: the average velocity over an interval equals a weighted combination of the average velocities over two sub-intervals split at an intermediate point c. Applying the MeanFlow identity to each sub-interval separately yields two derivative terms, one covering the interval from c to t (the early, stable phase) and one covering r to c (the late, volatile phase). Splitting the estimate this way also means the second evaluation at c anchors the estimate and limits how much early errors get amplified later.
They then study how to pick c: randomly, at the midpoint, or with a small learned network. Implementation-wise, the derivative is computed with a Jacobian-vector product using forward-mode automatic differentiation (dual numbers) in one forward pass, and because the two sub-interval terms are independent under stop-gradient, both are computed together by concatenating along the batch dimension. An extra embedding layer is added for the timestep r and summed with the t embedding.
Why This Matters
-
Research impact: The paper argues that one-step generation methods developed for image generation do not transfer to generalist RFMs, and backs this with a measurable property of the RFM velocity field rather than only benchmark numbers. It also reframes single-step action generation as a problem of estimating a time derivative under unequal noise conditions across the denoising trajectory.
-
Real-world applications:
- Real-time manipulation on robot-mounted edge hardware (for example NVIDIA Jetson Orin), where lower action-head latency translates directly into higher control frequency.
- Large-scale tabletop manipulation with GoogleX and WidowX embodiments, as evaluated through the Fractal and BridgeData V2 datasets in Real2Sim.
- Navigation policies on humanoid platforms (Unitree G1 with the COMPASS-generated point navigation dataset), including out-of-distribution settings.
- Deployment of pre-trained RFMs where a single-step action head can be substituted into an existing flow-matching model by fine-tuning, without retraining the whole system.
-
Industry relevance: The results are reported on GR00T-N1.6 and SimVLA-S, both generalist RFMs with transformer action heads, and the efficiency measurements span both a desktop GPU (L40) and an embedded robotics platform (Jetson Orin) in eager and compiled execution modes. Since action-head latency is reported as 43%–73% of end-to-end latency in the baseline, this is a direct lever on the decision rate of deployed robot policies.
Future Directions
- Reducing the training overhead. K-MF adds 7.5%–13.6% training time over MeanFlow and 33%/27% over flow matching on GR00T-N1.6 and SimVLA respectively; the authors list this as a limitation, leaving room for cheaper formulations.
- Extending to untested control domains. The paper states its evaluation does not cover whole-body control or dexterous manipulation, both of which it describes as still challenging.
- Understanding the dataset dependence. Spread in local acceleration correlates with training-data task diversity and dataset complexity, but the paper does not report a way to predict that spread before training — a potential direction for deciding when K-MF is needed.
- Improving the learnable intermediate point. The learned λ correlates with interval width and outperforms both random sampling and the fixed midpoint, raising the question of whether richer generators or explicit constraints on λ could do better.
Target Audience
- Robotics researchers working on generative action policies for manipulation and navigation.
- Machine learning researchers interested in few-step and one-step generation, flow matching, MeanFlow, and consistency-based methods.
- Engineers deploying robotic foundation models on latency-constrained edge hardware such as NVIDIA Jetson Orin.
- Graduate students and practitioners with background in diffusion/flow models who want a worked example of adapting an image-generation technique to robotics.
Authors’ abstract
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.