Skip to content
AI.info

Research

Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning

Overview Research area: Reinforcement learning — specifically offline-to-online RL, with generative models (flow matching) used as policies. Technical level: Advanced. The paper assumes familiarity wi

Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
arXiv
2602.18117
Published
2026-02-20
Authors
Yongjae Shin, Jongseong Chae, Jongeui Park, Youngchul Sung

AI summary

Overview

Research area: Reinforcement learning — specifically offline-to-online RL, with generative models (flow matching) used as policies.

Technical level: Advanced. The paper assumes familiarity with Markov Decision Processes, flow matching / continuous normalizing flows, ordinary differential equations, and standard offline RL algorithms such as CQL and IQL.

Scope: The paper proposes FINO, a method that injects controlled noise into a flow-matching policy during offline pre-training and then uses entropy-guided sampling during online fine-tuning, evaluated on 45 tasks from OGBench and D4RL under a limited online interaction budget.

What This Paper Is About

Generative policies built on flow matching perform well in offline RL, but when they are extended to an online fine-tuning phase, prior work simply continues training them rather than adapting anything for the online stage. The problem is that a flow-matching policy trained with the standard objective collapses onto the data points in the offline dataset, so it explores very little once the agent starts interacting with the environment — the authors illustrate this with an FQL agent on antmaze-giant-navigate that stays near the start and only ever reaches the goal via the upper route. The goal of this paper is to deliberately widen the action distribution the policy represents during offline pre-training, and then exploit that widened distribution to explore efficiently during a short online fine-tuning budget.

Key Contributions

  1. A noise-injected flow-matching objective for offline pre-training. Instead of matching the standard flow-matching target on the interpolated point x_t, FINO matches it on x_t + ε_t, where ε_t ~ N(0, α_t²I) has a scheduled variance α_t² = (η² − 2η)t² + 2ηt. The scheme reduces to standard flow matching when η = 0.

  2. Theoretical characterization of the induced distribution. Proposition 1 shows the injected noise induces conditional probability paths with the same mean t·x_i as flow matching but a variance (1 − (1 − η)t)² that is greater than or equal to flow matching's. Theorem 1 shows FINO still yields a valid continuous normalizing flow satisfying the continuity equation, and Theorem 2 shows the marginal probability path has variance greater than or equal to that of standard flow matching at every time t ∈ [0, 1] — larger variance at t = 1 meaning wider action coverage.

  3. An entropy-guided sampling mechanism for online fine-tuning. Rather than always taking the highest-value action among candidates, the agent draws actions from a Boltzmann distribution over Q values with temperature ξ, and adjusts ξ online via ξ_new = ξ − α_ξ[H − H̄] using policy entropy H and target entropy H̄.

  4. A large empirical study. Experiments on 45 tasks from OGBench and D4RL, with 1M offline pre-training steps followed by 500K online fine-tuning steps, 10 seeds, and 95% confidence intervals, plus ablations on the noise injection point, the entropy mechanism, and the action sample size.

Main Findings

  • Large gains in the hardest navigation tasks: On OGBench humanoidmaze-medium-navigate, FINO goes from 50 ± 7 after offline pre-training to 97 ± 1 after fine-tuning, versus FQL's 53 ± 6 → 61 ± 2. On humanoidmaze-large-navigate, FINO reaches 33 ± 7 versus FQL's 10 ± 3.

  • Strong D4RL results: On the D4RL antmaze aggregate over six tasks, FINO scores 79 ± 4 → 96 ± 1 (FQL 80 ± 4 → 95 ± 1; ReBRAC 80 ± 5 → 89 ± 5). On the D4RL adroit aggregate over four tasks, FINO reaches 112 ± 1 from 13 ± 2, the highest in the table.

  • No loss of offline performance: The authors state that gains are achieved without degrading offline performance, which they attribute to Theorem 2 — the induced path preserves the mean while increasing variance.

  • Action candidates alone are not the explanation: The comparison with IFQL, which also samples multiple actions per state, indicates that candidate sampling by itself does not account for the improvement.

  • Noise injection point matters: A baseline that injects Gaussian noise directly into actions (with the candidate mechanism and entropy guidance retained) underperforms. On door-cloned, where FQL's performance is approximately 100, adding noise to the action fails to help and can degrade performance.

  • Entropy-guided sampling beats entropy-scaled noise: An ER-Noise baseline that scales noise according to entropy is outperformed, particularly in high-dimensional action spaces such as humanoidmaze.

  • Both components are necessary: Removing noise injection (w/o Noise) or entropy guidance (w/o Guidance) degrades performance; the left plot of Figure 6 indicates noise injection is especially critical, while removing guidance hurts in later stages.

  • Efficiency is comparable: Training time increases slightly relative to the backbone due to entropy estimation and candidate sampling, but the increase is negligible compared to algorithms such as Cal-QL; inference requires fewer samples than baselines such as IFQL.

  • Sample size trade-off: Performance generally improves with more action candidates N_sample, with diminishing returns beyond a point; the authors set it to half the action dimension and use η = 0.1 for all environments since actions are bounded in [−1, 1].

Methodology in Plain English

The authors start from Flow Q-Learning (FQL), which represents a policy as a state-conditioned flow model trained to reproduce actions in the offline dataset, plus a one-step policy distilled from it for fast action selection.

Their change is in how the flow model is trained. Standard flow matching teaches the model the direction from a random noise sample to a dataset action, and because the conditional variance is set to zero, the model learns to place actions exactly on the data points. FINO instead adds a time-dependent amount of noise to the interpolated point that the model sees, and slightly rescales the target direction by a factor (1 − η). The variance schedule is designed so that no noise is added at t = 0 but variance η² remains at t = 1. The effect is that the model learns a slightly wider, but still data-centred, action region — shown visually in a toy experiment with a two-dimensional action space and a dataset placed inside four circular regions, where the flow-matching model hugs the data points while FINO's samples cover a wider area.

For the online stage, the agent samples several candidate actions per state using different noise inputs. It scores each candidate with the learned Q-function and converts the scores into a probability distribution using a temperature ξ, then samples an action from that distribution instead of always taking the argmax. A smaller ξ flattens the distribution toward exploration; a larger ξ sharpens it toward exploitation. Rather than fixing ξ, the method updates it toward a target entropy using the current estimated policy entropy, which is obtained by sampling multiple actions from the same state and fitting a Gaussian Mixture Model, following prior work. At inference time, the agent deterministically picks the highest-value action.

Why This Matters

Impact on research: The paper argues that offline-to-online RL should not be treated as a mere continuation of offline RL, and demonstrates that deliberately designing the offline pre-training phase to produce a broader action distribution pays off during fine-tuning. Its theoretical results connect an explicit training-time noise schedule to the variance of the resulting marginal flow distribution, giving a principled account of how a flow policy's support can be widened without shifting its mean. It also contributes the first use of a generative policy's expressivity directly for exploration in this framework, as opposed to prior work that used diffusion models for data augmentation.

Real-world applications (bullets):

  • Robotics: Learning manipulation skills such as cube-double-play and puzzle-4x4-play from a fixed set of demonstrations, then refining them with a short period of real robot interaction.
  • Navigation and logistics: Maze-navigation tasks such as antmaze and humanoidmaze translate to warehouse routing, autonomous vehicle path planning, and delivery routing where exploring alternative routes matters.
  • Industrial control: Settings where collecting large online datasets is expensive or risky, so a policy must be pre-trained on logs and adapted with a small interaction budget.
  • Simulation-to-reality transfer: Pre-training in simulation and fine-tuning briefly on a physical system, where sample efficiency is the binding constraint.

Industry relevance: Many industrial RL deployments have abundant logged data but very limited opportunity for live interaction. FINO's design target — strong performance under a limited online budget — matches that constraint, and the paper's efficiency analysis suggests the extra machinery does not add significant overhead relative to baselines.

Future Directions

  • The paper's Conclusion is truncated in the provided content, so the authors' own stated next steps are not fully reported. The following are open questions the work raises.

  • Choosing η in general action ranges: The authors fix η = 0.1 because all experimental environments use actions bounded in [−1, 1]. How to select η when action scales differ is not resolved.

  • Scaling N_sample: The paper finds diminishing returns beyond a certain number of candidate actions and adopts half the action dimension as a practical compromise; whether better selection schemes could get more out of fewer candidates is an open question.

  • Entropy estimation via Gaussian Mixture Models: Entropy is estimated by fitting a GMM to sampled actions because the distilled one-step policy's distribution is intractable. Whether more accurate or cheaper entropy estimates would improve the guidance is not reported.

  • Extending beyond flow matching: The noise-injection scheme is derived specifically for the flow-matching conditional probability path. Whether the same variance-preserving argument transfers to other generative policy classes is not addressed.

Target Audience

Researchers and graduate students working on reinforcement learning, generative models as policies, and offline-to-online learning will get the most from this paper. It is also relevant to practitioners who must adapt policies from logged data with a small live-interaction budget, and to readers interested in the theory of continuous normalizing flows, since the Proposition and Theorems give an explicit link between a training-time noise schedule and the variance of the learned marginal distribution. The level of the paper is advanced; readers should already be comfortable with flow matching, ODE-based generative models, and standard offline RL baselines.

Authors’ abstract

Generative models have recently demonstrated remarkable success across diverse domains, motivating their adoption as expressive policies in reinforcement learning (RL). While they have shown strong performance in offline RL, particularly where the target distribution is well defined, their extension to online fine-tuning has largely been treated as a direct continuation of offline pre-training, leaving key challenges unaddressed. In this paper, we propose Flow Matching with Injected Noise for Offline-to-Online RL (FINO), a novel method that leverages flow matching-based policies to enhance sample efficiency for offline-to-online RL. FINO facilitates effective exploration by injecting noise into policy training, thereby encouraging a broader range of actions beyond those observed in the offline dataset. In addition to exploration-enhanced flow policy training, we combine an entropy-guided sampling mechanism to balance exploration and exploitation, allowing the policy to adapt its behavior throughout online fine-tuning. Experiments across diverse, challenging tasks demonstrate that FINO consistently achieves superior performance under limited online budgets.

Read the original paper