Skip to content
AI.info

Research

Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control

Overview Research area: Robotics — robot learning, vision-language-action (VLA) policies, and generative (diffusion / Diffusion Transformer) pretraining for control. Technical level: Advanced. The pap

Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control
arXiv
2609.28339
Published
2026-09-23
Authors
Zanyi Wang, Yuheng Lei, Dengyang Jiang, Ping Luo, Mengdi Wang, Zhixuan Liang, Shilong Liu

AI summary

Overview

Research area: Robotics — robot learning, vision-language-action (VLA) policies, and generative (diffusion / Diffusion Transformer) pretraining for control.

Technical level: Advanced. The paper assumes familiarity with diffusion denoising trajectories, Diffusion Transformers, flow/velocity matching, mixture-of-transformers (MoT) architectures, and standard robot manipulation benchmarks.

Scope: A one-sentence summary: the paper shows that the future-prediction target used by current generative robot policies is not what makes generative co-training work, and that applying the pretrained denoising process directly to the current visual stream used for control (a method called NowWAM) improves robustness while halving training visual tokens.

What This Paper Is About

Recent robot policies build on pretrained generative Diffusion Transformers (DiTs) by co-training them alongside an action predictor, most often by asking the model to predict future visual observations. Over time, these methods have moved future prediction out of deployment and kept it only as a training-time target, which raises a basic question: if the future is never used at inference, is predicting the future actually the reason co-training helps?

The paper investigates what a pretrained generative DiT contributes to action learning, and finds that the benefit is not tied to forward temporal semantics. The authors propose NowWAM, a formulation that removes the separate future target entirely and instead denoises the same current visual stream that the action expert consumes, coupling generative supervision to action learning across the denoising trajectory.

Key Contributions

  1. A dissection of how generative vision priors transfer to control. Controlled analyses show that past and future visual targets perform comparably (83.80% for the past target at t−16 versus 82.89% at t+16 and 83.70% at t+32), while restricting adaptation to the clean endpoint of the denoising process substantially reduces robustness (77.48%).
  2. NowWAM, a future-target-free generative co-training policy. It applies current-frame denoising to the action-facing representation, coupling generative adaptation and action learning in a single visual stream, with the action pathway left unchanged.
  3. Evidence that this simpler interface generalizes across benchmarks and architectures. NowWAM improves performance on LIBERO-Plus (87.7% with FLUX2-Klein, 87.8% with the pure text-to-image Z-Image backbone) and RoboCasa GR1 Tabletop (64.9% with 100 demonstrations per task), and halves the visual tokens used during joint training.
  4. A training-efficiency comparison. Removing the second visual stream cuts step time from 2.85 s to 1.63 s (1.8× faster) and peak memory from 78.9 GiB to 74.3 GiB under the same 2×H200 setup with global batch size 64.

Main Findings

  • Future prediction is not uniquely responsible for the benefit. Under strictly matched controlled settings, a current target (79.10%), future targets at t+16 (82.89%) and t+32 (83.70%), and a past target at t−16 (83.80%) all perform in a similar range, with the past target slightly ahead of both future targets.
  • The denoising trajectory is the important interface. Training only at the clean endpoint (σ=0) with the visual objective retained reaches 77.48%, far below uniform denoising-state sampling (83.46%) and low-noise-shifted sampling (84.97%). Removing the visual denoising loss but keeping noisy-current exposure reaches 78.73%.
  • The pretrained generative backbone is a major robustness prior on its own. Pretrained initialization with the Edit objective raises success from 58.14% to 83.05% over random initialization; even without any visual objective, pretraining improves performance from 50.55% to 75.81%.
  • LIBERO-Plus robustness improves over the future-target baseline by 6.1 points. NowWAM with FLUX2-Klein reaches 87.7% overall versus 81.6% for ImageWAM (FLUX2-Klein-4B), with gains concentrated in camera (88.6 vs 77.7), robot (70.5 vs 48.3), background (93.0 vs 86.0), and noise (97.4 vs 95.2) perturbations; language performance is comparable (87.5 vs 89.9).
  • The formulation works with a pure text-to-image backbone. NowWAM with Z-Image-6B reaches 87.8% on LIBERO-Plus, slightly above the FLUX2-Klein result of 87.7%, showing the approach does not depend on video generation or image-editing pretraining.
  • In-distribution manipulation is preserved. On standard LIBERO, NowWAM averages 98.4% (Spatial 99.5, Object 100.0, Goal 97.5, Long 96.5), matching the future-target generative baseline (ImageWAM at 98.4%) in a near-saturated setting.
  • Few-shot cross-benchmark transfer improves. On RoboCasa GR1 Tabletop with only 100 target-task demonstrations, NowWAM reaches 64.9%, above DIAL (58.3%) and Fast-WAM (56.0%) at the same 100-shot budget, over 1,200 evaluation episodes.
  • Training cost drops with no added acceleration. Removing the second visual stream reduces visual tokens from 784 to 392 (50% fewer), step time from 2.85 s to 1.63 s (1.8× speedup), and peak memory from 78.9 GiB to 74.3 GiB (5.8% lower) under identical hardware and batch size.

Methodology in Plain English

The authors start from a standard generative co-training recipe: a pretrained image/video Diffusion Transformer processes the current observation plus a separate target image (usually a future frame), and the target is noised along the pretrained denoising trajectory while a separate action expert learns to predict robot actions. The target branch is dropped at deployment.

NowWAM changes only where the generative objective is applied. Instead of adding a separate future image, the model takes the current observation's latent, adds noise to it along the same denoising trajectory the pretrained DiT already knows, and asks the backbone to predict the corresponding velocity (the noise minus the clean latent, ε − z_t). The same noised representation is what the action expert reads, so the action loss and the generative loss are computed on one shared visual stream rather than two. The noise level is drawn from a distribution that can be tuned toward the low-noise end, and the total loss is a weighted sum of the denoising loss and a masked mean-squared action loss, with λ_vis = 0.5 and λ_act = 1.0. At inference, the policy simply evaluates the clean endpoint (σ = 0) in a single forward pass — no future rollout, no separate visual stream.

Architecturally, the policy keeps a DiT–MoT structure: a pretrained generative DiT backbone and a separate action DiT whose tokens and parameters remain distinct, interacting through masked mixed attention. The evaluation uses standard LIBERO as an in-distribution reference, RoboCasa GR1 Tabletop for few-shot transfer, and LIBERO-Plus as the primary robustness benchmark across seven perturbation categories (10,030 episodes under the full protocol). Controlled ablations, which vary only the stated intervention, use non-EMA checkpoints on a fixed 1,923-episode LIBERO-Plus subset, while main results use final EMA checkpoints.

Why This Matters

Impact on research. The paper challenges an assumption baked into much recent generative-policy work — that the future target is the mechanism by which generative pretraining transfers to control. By showing past and future targets perform comparably, and that denoising-state exposure matters more than the endpoint, it reframes generative adaptation as being about the native denoising interface rather than forward temporal prediction. It also shows the recipe transfers to a pure text-to-image DiT, which broadens the pool of usable backbones.

Real-world applications (as motivated by the paper's settings):

  • Robotic manipulation in visually messy or shifted environments (camera movement, robot appearance changes, background clutter, sensor noise), where the robustness gains on LIBERO-Plus are concentrated.
  • Few-shot adaptation to new simulators and task families, per the RoboCasa GR1 Tabletop 100-demonstration setting.
  • Cost-constrained policy training pipelines, since the single-stream formulation reduces visual tokens and step time without extra acceleration.
  • Deploying policies from image-generation backbones that were never trained on video or image editing, demonstrated with Z-Image-6B.

Industry relevance. Halving training visual tokens and cutting step time from 2.85 s to 1.63 s while lowering peak memory from 78.9 to 74.3 GiB directly reduces the cost of fine-tuning large generative backbones for robot control, and a simpler single-stream interface lowers engineering complexity for teams adapting generative models to action prediction.

Future Directions

  • When does explicit future modeling still pay off? The authors state that future modeling may still be useful for explicit dynamics or long-horizon planning, but is not required in the manipulation settings studied; identifying which task regimes need it remains open.
  • How far does the low-noise shift generalize? Shifting the sampling distribution toward low noise improved success from 83.46% to 84.97% in the controlled subset; whether this optimum varies by backbone, task, or dataset is not established.
  • Does the formulation scale beyond the studied backbones and benchmarks? The paper tests FLUX.2-Klein-4B and Z-Image-6B on LIBERO, LIBERO-Plus, and RoboCasa GR1 Tabletop; behavior with other generative backbones, real-robot data, or other embodiments is not reported.
  • Can the robustness prior from pretrained initialization be isolated further? Pretrained initialization improved success even with no visual objective (75.81% vs 50.55%), suggesting an unexplored component of the benefit that is separate from the generative co-training objective.

Target Audience

Robotics and embodied-AI researchers working on vision-language-action models, diffusion policies, and world-model-style co-training; practitioners adapting large generative image or video models to control; and engineers who need to reduce the fine-tuning cost of generative backbones for manipulation, whether in simulation or on physical robots. Readers should be comfortable with diffusion/flow-matching notation and standard manipulation benchmark protocols.

Authors’ abstract

Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.

Read the original paper