Skip to content
AI.info

Research

Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models

Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models Authors: Lior Cohen (Technion), Ofir Nabati (Technion), Kaixin Wang (Microsoft Research), Navdeep Kumar (Technion), Shie Mann

Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
arXiv
2602.08032
Published
2026-02-08
Authors
Lior Cohen, Ofir Nabati, Kaixin Wang, Navdeep Kumar, Shie Mannor

AI summary

Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models

Authors: Lior Cohen (Technion), Ofir Nabati (Technion), Kaixin Wang (Microsoft Research), Navdeep Kumar (Technion), Shie Mannor (Technion) arXiv: 2602.08032v2 [cs.LG], 17 Feb 2026 — License: CC BY 4.0

Overview

Research area: Reinforcement learning with diffusion-based world models, specifically the efficiency of imagination-based policy training for discrete stochastic policies.

Technical level: Advanced. The paper assumes familiarity with diffusion/rectified flow formulations, POMDPs, actor-critic methods, denoising schedules, and latent world model architectures.

Scope: The paper proposes Horizon Imagination (HI), a method for generating multiple future observations in parallel during world model imagination, together with a stable discrete action sampler and a new sampling schedule, evaluated on Atari 100K and Craftium.

What This Paper Is About

Diffusion world models generate highly realistic observations, but using them to train controllers is expensive: each imagined observation requires a multi-step denoising process, and imagination normally interleaves policy decisions with world model predictions one step at a time, making the whole rollout inherently sequential. The authors' goal is to make on-policy imagination in diffusion world models fast enough to be practical by denoising many future observations at once, without losing control performance or generation quality.

Key Contributions

  1. Horizon Imagination (HI): an on-policy imagination procedure for discrete stochastic policies that denoises multiple future observations in parallel, reducing the sequential burden of diffusion generation. The approach is described as training-agnostic and applicable to any pre-trained world model with observation-level time conditioning.

  2. A stable discrete action sampling mechanism: a sampling scheme based on a single multivariate uniform sample and a fixed random permutation over actions, which guarantees consistent action selection across evolving policy distributions. The authors prove a proposition showing the scheme is unbiased for any categorical distribution and that the probability of an action change between two distributions is bounded below by their total variation distance and above by the L1 distance between the corresponding conditional-probability vectors.

  3. The Horizon schedule: a sampling schedule defined as a matrix K of denoising times that disentangles the denoising budget B from the decay horizon ν, allowing any combination of the two and, critically, sub-frame budgets where B < h (fewer denoising steps than observations). Standard autoregressive generation requires B ≥ h and is limited to multiples of h.

  4. Empirical validation and analysis: control benchmarks on four Atari 100K games and four Craftium games, plus an FVD/MSE study of generation quality across ν and B spanning fully autoregressive to highly parallel regimes, along with an ablation of the stable sampling mechanism and a comparison against the Pyramidal schedule of Chen et al. (2024).

Main Findings

  • Sub-frame budgets preserve control performance: Baselines with decay horizon ν = 4 maintained the performance of the autoregressive baseline (ν = 1, B = 32) across all environments at reduced cost. Full performance was sustained with a sub-frame budget of B = 16, i.e. only half the denoising steps of the autoregressive baseline.

  • One denoising step per observation generally suffices: Results suggest a single denoising step per observation is generally enough, with B = 32 baselines performing comparably across environments. The exception reported is Craftium/ChopTree-v0, the most visually complex environment, where the (ν = 4, B = 32) baseline achieved superior performance.

  • Atari is visually simpler for generation: The authors report that in Atari even a single denoising step could yield satisfactory generation quality, whereas Craftium presents richer, more complex observations and is a harder test for the world model.

  • Stable sampling is critical: Ablating stable action sampling in favor of drawing a fresh action before each denoising step caused a substantial drop in returns; a particularly prominent collapse was observed on Atari Boxing and Gopher.

  • Action-change counts near the theoretical floor: In a controlled experiment with N actions and 10^6 sampled (ω, ρ) pairs per distribution pair over 10^4 Dirichlet-sampled pairs, the method performed very close to the total-variation lower bound across all N values. In a simulated 16-step denoising process with N = 10 actions and 1,000 source–target pairs per entropy setting, the proposed method yielded at most one action change on average across all 16 steps, whereas naive sampling altered more than half overall. For high-entropy distributions, action changes decreased under the proposed method but increased under naive sampling.

  • Parallel generation is advantageous for quality: FVD results showed a consistent trend favoring parallel configurations (4 ≤ ν ≤ 16) under low to medium budgets. In several cases, parallel variants with sub-frame budgets achieved quality comparable to baselines using a budget 16 times larger.

  • Extreme budgets favor autoregressive: At B ≥ 128, performance tended to degrade slightly as ν increased, with the autoregressive baseline consistently ranking at the top.

  • Perceptual quality and ground-truth fidelity diverge: MSE results suggested that increasing the denoising budget may cause generated sequences to drift further from the ground truth even as perceptual (FVD) quality improves.

  • Pyramidal schedule collapses at higher budgets: Under the Pyramidal schedule of Chen et al. (2024), whose decay horizon drifts with budget because the two are entangled, the authors report a pronounced collapse in performance as budgets grow (detailed in Appendix D).

  • Runtime reduction: On Atari, training with B = 32 required approximately 27 hours per run, while B = 16 shortened training to about 19 hours. Because tokenizer and world model training stages are unaffected by the denoising budget, the overall speed-up is smaller than a full 2×.

Methodology in Plain English

The agent has four parts: an image tokenizer that compresses observations into latents constrained to [-1, 1] via a tanh activation, a diffusion-based world model (a causal Diffusion Transformer conditioned on actions and denoising time) that predicts future latents, a lightweight reward–termination predictor, and an actor-critic controller. The world model is trained on short trajectory segments sampled from a replay buffer, with an independent denoising time sampled for each observation, as in Diffusion Forcing (Chen et al., 2024); in 20% of cases a clean prefix is supplied by setting the denoising times of the earliest frames to 1, with the prefix length drawn uniformly from {1, …, ⌊0.7h⌋}.

The distinctive step is at inference. Rather than fully denoising one observation before moving to the next, HI denoises all h future observations simultaneously. Because later observations need actions, and those actions must be conditioned on still-noisy observations, the policy is queried before every denoising step, producing an increasingly informed action distribution for each timestep.

Two problems arise and both are addressed directly. First, naively resampling actions at each denoising step causes frequent action flips that destabilize generation. The fix: at the start, draw a single uniform vector and a random permutation over the action set, then map those fixed random objects through the evolving distribution at every step. The same underlying randomness therefore yields consistent choices, and identical distributions yield identical actions. Second, the schedule: the Horizon schedule defines, for each denoising step b and each future timestep t, a denoising time along a line with slope −1/ν, clipped to [0, 1], so that near-future observations are denoised before far-future ones, and the decay horizon ν can be set independently of the total budget B.

The controller is trained actor-critic style. The critic sees only fully denoised inputs; the actor is trained with REINFORCE at every trajectory step before termination, with an entropy term, but updates are restricted to steps where the denoising time of the next observation increases, to balance updates across denoising times.

Why This Matters

Impact on research: The paper reframes an efficiency question that the diffusion world model literature has largely sidestepped: existing large diffusion world models are motivated by agent training but are limited to conditional video generation and do not address control, while the methods that do address control either pay heavy sequential costs or are restricted to continuous action spaces. HI offers an on-policy imagination route for discrete actions and provides a systematic study, previously missing, of how generation quality depends on sequential versus parallel sampling and on the denoising budget.

Real-world applications (contexts implied by the authors' stated motivation for lightweight, efficient controllers):

  • Deployment of learned controllers on hardware where real-time inference and low power consumption are required.
  • Domains where repeatedly querying a large world model at controller inference time is impractical.
  • Training pipelines that rely on large-scale simulated experience to reduce costly real environment interaction.
  • Environments with complex visual observations and discrete action spaces, such as Craftium-style 3D settings.

Industry relevance: The method is training-agnostic and works with any pre-trained world model that uses observation-level time conditioning, which means the schedule and stable sampler can be dropped into existing diffusion world model stacks rather than requiring retraining from scratch. The authors' design also separates a large general-purpose dynamics module from small, task-specific reward–termination predictors, so new tasks require learning only the small predictors.

Future Directions

  • Broader schedule sweeps: The authors explicitly note that although ν = 4 consistently stood out, they did not extend the study to additional configurations, because evaluating each baseline requires full runs across all environments with five seeds. Future work could explore more configurations.

  • Scaling beyond eight environments: The study was restricted to eight environments to keep evaluations tractable; broader or more diverse benchmark coverage remains open.

  • Reconciling FVD and MSE: The observed divergence, where larger budgets improve perceptual quality but move generated sequences further from ground truth, is reported but not resolved.

  • Reattacking extreme budgets: Performance at B ≥ 128 favors the autoregressive baseline and degrades slightly as ν increases, leaving open how to make highly parallel generation competitive at very large budgets.

Target Audience

This paper is most valuable to reinforcement learning researchers and engineers working on model-based agents, diffusion-based generative models, or video world models, particularly those concerned with the compute cost of imagination-based training and with deploying lightweight policies. It will also interest practitioners building on Diffusion Forcing-style or Pyramidal-schedule world models who want to reduce denoising budgets without losing control performance. Readers need a working knowledge of diffusion sampling, actor-critic RL, and latent world model pipelines; the paper is not an introductory treatment.

Authors’ abstract

We study diffusion-based world models for reinforcement learning, which offer high generative fidelity but face critical efficiency challenges in control. Current methods either require heavyweight models at inference or rely on highly sequential imagination, both of which impose prohibitive computational costs. We propose Horizon Imagination (HI), an on-policy imagination process for discrete stochastic policies that denoises multiple future observations in parallel. HI incorporates a stabilization mechanism and a novel sampling schedule that decouples the denoising budget from the effective horizon over which denoising is applied while also supporting sub-frame budgets. Experiments on Atari 100K and Craftium show that our approach maintains control performance with a sub-frame budget of half the denoising steps and achieves superior generation quality under varied schedules. Code is available at https://github.com/leor-c/horizon-imagination.

Read the original paper