Skip to content
AI.info

Research

Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

Variance-Averse n-Step Offline RL for Sparse Long-Horizon Environments Overview Research area: Offline reinforcement learning (RL), specifically generative (flow-matching) policies, distributional val

Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments
arXiv
2610.07899
Published
2026-10-06
Authors
Guhyeon Kang, Minhae Kwon

AI summary

Variance-Averse n-Step Offline RL for Sparse Long-Horizon Environments

Overview

Research area: Offline reinforcement learning (RL), specifically generative (flow-matching) policies, distributional value estimation, and risk-sensitive action selection.

Technical level: Advanced. The paper assumes familiarity with Markov decision processes, actor–critic methods, n-step Bellman targets, categorical distributional critics (C51-style), flow matching, and convex-order stochastic dominance.

Scope: The paper introduces VAN-Flow, an offline RL framework that combines a categorical distributional critic, a new "variance-averse expectation" operator over return distributions, and a flow-matching actor with rejection sampling, and evaluates it on more than 40 tasks from D4RL and OGBench.

What This Paper Is About

Generative actors in offline RL can reproduce every mode present in a heterogeneous dataset — including action modes that occasionally produce high returns by chance but are not dependable. Because a standard critic only exposes the mean return (the expected Q-value), maximizing it alone cannot distinguish a reliable action from a lucky one, and this problem is amplified by n-step returns, which add variance. The paper's goal is to steer a flow-matching actor toward actions whose return distributions are simultaneously high-return and low-variance, using a single aggregation operator rather than a new risk-sensitive objective or penalty term.

Key Contributions

  1. Diagnosis of a generative-actor failure mode. The authors formalize "reliable behavior" as low dispersion of the dataset return distribution and show empirically that expected Q-values alone are insufficient to guide generative policies under heterogeneous data. Their Figure 1 shows this on antmaze-large in low-variance (navigate) and high-variance (explore) regimes, where the flow-based method FQL-n degrades far more than the Gaussian ReBRAC-n under the same n-step setting.

  2. The variance-averse expectation ℰ(Z). A new operator over categorical return distributions that smoothly reweights atom probabilities to favor high-return and low-variance actions. It has no hard truncation and no auxiliary penalty term, and reduces to the ordinary expectation when δ = 0. Theorem 3.2 proves convex-order monotonicity of the underlying spectral form and bounds the discretization gap of the implemented operator.

  3. The VAN-Flow framework. A unified actor–critic method combining (i) a categorical distributional critic trained with an n-step distributional Bellman target, (ii) ℰ(·)-guided rejection sampling over M candidate actions, and (iii) a flow-matching actor trained with a single Euler step as a critic-gradient surrogate.

  4. Broad empirical validation. More than 40 tasks from D4RL (AntMaze) and OGBench, with the largest gains in long-horizon and high-variance regimes, plus ablations, sensitivity analyses, a runtime comparison, and an off-to-online fine-tuning study (Appendix J).

The authors state this is, to their knowledge, the first work to use variance-averse aggregation over categorical return distributions to select reliable actions for generative actors in offline RL, and to show empirically that this matters for stable n-step learning under heterogeneous data.

Main Findings

  • Stability of the operator is certified. Theorem 3.2 shows the spectral form ℰ^sp prefers the less dispersed of two equal-mean distributions (convex-order monotonicity), and bounds the gap between it and the implemented ℰ by R·ε_δ(Z), where R = z_I − z_1 and ε_δ(Z) ∈ [0, 1]. On trained critics the realized gap is a median 0.37% of R, and ℰ and ℰ^sp select the same rejection-sampling candidate in 96.9% of states with Kendall's τ = 0.99.

  • The gap shrinks with more atoms. If Z is the C51 projection of a target supported on [V_min, V_max] with Lipschitz-continuous CDF, then max_i p_i = O(1/I) and ε_δ(Z) = O(1/I), so the gap vanishes as the number of atoms I grows.

  • VAN-Flow leads on OGBench. Among the reported OGBench results, VAN-Flow scores 95 on antmaze-large-navigate, 81 on humanoidmaze-large-navigate, 92 on humanoidmaze-giant-navigate, 68 on antmaze-giant-navigate, and 84 on the high-variance antmaze-large-explore task, versus 58 for QC, 48 for ReBRAC, and 1 for FQL-n on that last task.

  • Flow-based n-step methods collapse on high-variance data. On antmaze-large-explore, FQL-n and BFN-n degrade much more than the Gaussian ReBRAC-n under the same n-step setting, while VAN-Flow remains robust — the paper's central motivating observation.

  • Strong D4RL AntMaze results. VAN-Flow reaches 95.6 ± 2.6 on umaze-diverse (IQL: 54.2 ± 5.5; ReBRAC: 83.5 ± 7.0) and 88.6 ± 3.0 on large-diverse (IQL: 30.2 ± 3.6; FQL: 83.0 ± 4.0; LEQ: 60.2 ± 18.3).

  • ℰ(Z) beats other aggregation rules. In Table 3, ℰ(Z) achieves the lowest normalized Var[Z] on all three tested tasks and the highest success: on humanoidmaze-large-navigate 0.93 variance / 87 success (vs. 77 for E[Z], 83 for CVaR, 79 for entropic, 80 for mean–variance); on antmaze-teleport-navigate 0.62 / 71; on antmaze-large-explore 0.66 / 93.

  • A single fixed hyperparameter works. δ = 2 is used across nearly all tasks, whereas CVaR, entropic risk, and mean–variance require per-task tuning of a confidence level or penalty coefficient β.

  • Both components of the critic/actor matter. Figure 3 shows Flow+Categorical outperforms partial combinations (flow with a regression critic, or a Gaussian actor), and the flow-based actor outperforms its Gaussian counterpart on long-horizon, sparse-reward problems.

  • Flow matching is more step-efficient than diffusion in their comparison. Appendix G reports flow matching achieving over 90% success rate with only 3 flow steps, while a diffusion-based variant requires 20 steps for comparable performance.

  • Rejection sampling and flow steps saturate quickly. Performance improves with more candidate actions but most tasks are stable using fewer than 32 candidates; a single integration step is insufficient, but performance is already high at three to five flow steps.

  • Q-guidance improves efficiency, not just success. The Q-guided actor achieves both higher success and higher time efficiency TE = 100·(1 − T/T_max); actors without Q guidance can reach high success on some tasks but with extremely low time efficiency.

  • Runtime is partly reported. Table 4 lists FQL at O(UKd²), 1.1 ms per step, 0.32 h wall-clock; BFN at O(UMKd²), 2.1 ms, 0.60 h; and QC at O(UMKd²), 3.0 ms, 0.83 h on humanoidmaze-large-navigate. The VAN-Flow row of that table is cut off in the provided content, so its runtime and complexity are not reported here.

  • n-step horizon shows a trade-off. Increasing n substantially improves long-horizon performance through better credit assignment, but excessively large n yields diminishing or negative returns due to amplified variance; VAN-Flow stays robust across a wide range.

Methodology in Plain English

The method rests on a simple idea: an action is "reliable" if its returns are consistent across the dataset, not just high on average. To measure consistency, the authors replace the usual single-number critic with a categorical critic that predicts a full return distribution over a fixed set of atoms, so both the mean and the spread are visible.

They then define the variance-averse expectation ℰ(Z), which reweights each atom's probability by a factor based on how far up the CDF that atom sits, raised to a power δ. Atoms in the lower part of the distribution gain relative weight; the more spread out the distribution, the more the score is dragged down. Unlike CVaR, which chops off the tail, or mean–variance and entropic risk, which bolt on a penalty term to the objective, this is a single reweighting step applied to the critic's output, with no extra hyperparameter to tune per task.

At training time the actor is a flow-matching policy: it learns a velocity field that transports noise into actions. To avoid expensive full ODE integration, the authors advance the flow with a single Euler step and evaluate the critic at that one-step action proxy, using it to pass a gradient back to the velocity field. Action selection uses rejection sampling: the actor proposes M candidate actions, and the one maximizing ℰ(Z(s, a)) is chosen. That same chosen action — not a freshly sampled actor action — is used as the target action in the n-step distributional Bellman backup, so value propagation is aligned with reliable-action selection. Standard target networks and double-Q are used for stability.

Evaluation covers D4RL AntMaze tasks and state-based OGBench tasks with noisy and suboptimal datasets, averaged over five random seeds, on an AMD Ryzen Threadripper PRO 9975WX CPU and an NVIDIA RTX PRO 6000 GPU.

Why This Matters

The paper targets a structural weakness of generative policies in offline RL: expressiveness without a way to tell dependable behavior from lucky behavior. It argues that this weakness is precisely what makes n-step returns fragile on heterogeneous data, and shows that a low-cost reweighting of a distributional critic's output mitigates it. For the broader field, it offers an alternative to the risk-sensitive-RL toolkit (CVaR, entropic risk, mean–variance) that requires no new objective and plugs directly into existing categorical critics.

The paper explicitly frames its motivation around robotics and physical AI foundation models built on generative actors; it does not evaluate those deployments.

  • Robotics and physical AI: the paper names generative actors as the backbone of foundation models for robotics and physical AI, where reliable action selection among multimodal candidates matters.
  • Offline training from mixed-quality logs (implied): the setting is heterogeneous data collected from behavior policies of varying quality — the situation in which reliability filtering is most valuable. The paper does not run experiments outside offline RL benchmarks.
  • Long-horizon planning with sparse rewards (implied): the benchmark emphasis on sparse, long-horizon tasks suggests relevance to tasks where successful trajectories are rare, though no non-benchmark domain is tested.
  • Risk-aware decision systems (implied): the operator produces a distribution-aware score, which could inform downstream selection, but the paper studies it only as an action-selection rule within VAN-Flow.

Industry relevance: the method is designed to work with standard categorical distributional critics and requires no modification to the underlying distributional RL training procedure, which lowers the barrier to adoption in existing offline RL pipelines. Rejection sampling contributing no extra penalty term and being stable with fewer than 32 candidates for most tasks suggests a modest inference overhead, though the paper's own runtime table for VAN-Flow is not available in the provided content.

Future Directions

  • Extending the reliability notion beyond one operator. The paper compares ℰ(Z) with CPW, Norm, Wang, Sharpe, and Sortino measures in Appendix K; how these alternatives fare systematically across the full 40-plus-task suite is an open comparison.
  • Tuning the aversion coefficient. δ = 2 is used for nearly all tasks; whether a principled, task-adaptive rule for δ exists — or whether its insensitivity is fundamental — is not resolved.
  • Scaling the atom count. Theorem 3.2 predicts the discretization gap vanishes as the number of atoms I grows; the empirical effect of larger I on critic cost and on action-selection agreement (currently 96.9% with τ = 0.99) is not reported.
  • Beyond the offline-to-online fine-tuning study. Off-to-online results appear in Appendix J, but how variance-averse selection interacts with online exploration and with continual data collection remains an open question.

Target Audience

Offline RL researchers working on distributional value estimation, risk-sensitive objectives, or generative (flow/diffusion) policies will get the most from this paper, particularly those dealing with sparse-reward, long-horizon tasks and heterogeneous datasets. Practitioners building offline RL pipelines on top of categorical critics or flow-matching actors will find the operator and the rejection-sampling scheme directly implementable. Readers without grounding in n-step returns, distributional RL, and stochastic-dominance concepts will find the theory section (Definition 3.1 and Theorem 3.2) demanding.

Authors’ abstract

Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.

Read the original paper