Skip to content
AI.info

Research

Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning

Overview Research area: Reinforcement learning — specifically directed exploration via optimistic value estimation and ensembles. Technical level: Intermediate. The paper assumes familiarity with Mark

Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
arXiv
2602.12375
Published
2026-02-12
Authors
Abdul Wahab, Raksha Kumaraswamy, Martha White

AI summary

Overview

Research area: Reinforcement learning — specifically directed exploration via optimistic value estimation and ensembles.

Technical level: Intermediate. The paper assumes familiarity with Markov Decision Processes, temporal-difference (TD) learning, DQN/Double DQN, and the idea of reward bonuses versus value bonuses, though the core idea (train an ensemble on random targets and use the error as a bonus) is describable in plain terms.

Scope: The paper introduces and analyzes VBE (Value Bonuses with Ensemble Errors), a drop-in exploration mechanism that adds an ensemble-derived value bonus to any base value-based RL algorithm, supported by two propositions, one high-probability guarantee, and experiments on four classic exploration environments plus six Atari environments.

What This Paper Is About

Directed exploration via optimistic value estimates has a practical adoption problem: reward-bonus methods (RND, ACB) only reward states retroactively, after the agent has already seen a high bonus there, so they cannot promote first-visit optimism. Bootstrap DQN (BDQN) does provide first-visit optimism, but requires replacing the learning algorithm with an ensemble-based one and choosing tricky hyperparameters like how often to resample a value function. The goal of this paper is a simple, easy-to-use exploration method that directly estimates a value bonus — a quantity added to the action-value only at action-selection time — with first-visit optimism, deep exploration, a small computational footprint, and compatibility with any base RL algorithm.

Key Contributions

  1. A new algorithm, VBE (Value Bonuses with Ensemble Errors). VBE maintains an ensemble of random action-value functions (RQFs). The bonus is the maximum absolute error between each RQF target and its learned predictor, and the behavior policy is greedy on q_w(s,a) + c·b(s,a). Because the RQF target comes from the same function class as the predictor, the error can go to zero, letting the bonus decay.

  2. Theory characterizing the bonuses. Proposition 1 proves that the action-values of the random reward induced by an RQF equal the RQF itself (q_i^π = f_i), which is the mechanism that lets bonuses shrink to zero. Proposition 2 shows the bonus can reflect MDP-specific transition stochasticity. Proposition 3 gives a condition on the bonus scale c under which initial values are optimistic with probability 1 − δ.

  3. A comparison to widely used approaches. The paper contrasts VBE with RND — whose bonuses come from random functions, not random value functions, and therefore miss MDP-specific properties — and with BDQN, showing (for a fixed policy) that BDQN's random-prior bonus is a scaled negation of VBE's ensemble reward, but that BDQN scales its targets by c, increasing target variance by a factor of c², while VBE uses c only in the behavior policy.

  4. An extensive empirical evaluation. Experiments span four classic exploration environments and six Atari environments, with a dedicated state-coverage study in Deepsea.

Main Findings

  • Reward-bonus methods lack first-visit optimism. The paper states that reward bonus approaches "do not encourage first-visit optimism" — a bonus can only rise retroactively, after the agent has already seen a high reward bonus from that state and action, so they cannot encourage visiting a state for the first time.

  • Bonuses can decay to zero. Proposition 1 shows q_i^π = f_i for all i ∈ [k], so updating the predictor with the ensemble reward should converge to the RQF target, which drives the value bonus to zero. The paper notes this convergence is only guaranteed under certain conditions (local minima, divergence due to TD with neural networks, off-policy updating). The paper reports that in its own experiments, the value bonuses always converged to zero, but states that no theory guarantees this when the behavior and target policies are both changing.

  • Bonuses capture MDP-specific properties. Proposition 2 expresses the bonus as a maximum over an expectation of a discounted difference of next-state values plus the Bellman error, which the paper uses to argue that VBE's bonuses can reflect the stochasticity in the transition dynamics of the MDP, unlike RND's.

  • Optimistic initial values with high probability. Proposition 3 gives an explicit lower bound on c: if c ≥ sqrt(n/π)(q_max − z(δ/2)/sqrt(n)) / (log(k/2) − log log(2/δ)), then q(s,a) + c·max_i |f_i(s,a) − f_{w_i}(s,a)| > q_max with probability 1 − δ, for any state and action.

  • Full state coverage in Deepsea. VBE covers the entire state space, even for the larger grid sizes tested; the paper reports that for a grid size of 50 the environment has 1275 total unique states, marked as a dotted line that VBE's unique-state progression reaches. The paper states VBE "consistently explores new states at a significantly higher rate" than the compared methods.

  • Better performance than the baselines on classic environments. The abstract states VBE outperforms Bootstrap DQN and two reward bonus approaches (RND and ACB) on several classic environments used to test exploration. The paper does not report tabulated numeric returns in the content available.

  • Scales to Atari. The paper provides "demonstrative experiments" that VBE can scale easily to more complex environments like Atari, without the design choices required to alter the underlying algorithm.

  • Small computational footprint. On each step only one RQF predictor is updated, which the paper states makes the per-step computation simply double that of Double DQN, regardless of ensemble size. Updating each predictor less frequently also makes the bonus decay more slowly, allowing a smaller ensemble.

  • PPO-based baselines were less sample efficient. Comparing against the released variants of ACB and RND that use PPO, the paper reports that the PPO version is generally less sample efficient than the DDQN versions.

  • Experiment scale. Classic-environment results use 50000 steps and 30 runs, except Deepsea, which uses 10000 episodes and 5 runs. Metrics differ by environment: accumulated reward over learning for River Swim (continuing), undiscounted episodic return for Deepsea and Puddle World, and discounted return for Mountain Car.

Methodology in Plain English

Instead of rewarding the agent for surprising outcomes and then learning a value function on those rewards, VBE separates the exploration signal from the main value function entirely.

The recipe is:

  1. Draw k random action-value functions (RQFs) — for example random neural networks — and never change them experimentally; they are fixed targets.
  2. For each RQF, define a random reward for a transition as the RQF's own value at the current state-action minus the discounted RQF value at the next state-action. Proposition 1 shows that the true value of this reward is exactly the RQF, so the target is representable.
  3. Maintain a learned predictor for each RQF and train it with the same bootstrapping machinery as the base algorithm. Since the target is representable, the error can shrink to zero.
  4. On each step, update only one randomly chosen predictor, keeping computation at roughly double Double DQN.
  5. Define the value bonus as the largest absolute difference between an RQF target and its predictor, and take the greedy action in q_w(s,a) + c·b(s,a), with c a scale parameter.

The ensemble value functions are updated on the same target policy as Double DQN — the greedy policy in q_w — because the paper wants to quantify uncertainty in the values of the target policy. Experiments compare this against DDQN-based variants of ACB and RND, DQN with additive priors (DQN-P, effectively BDQN with one value function in the ensemble), and BDQN, across Sparse Mountain Car, Puddle World, River Swim, and Deepsea, and then six Atari environments. In River Swim, the paper flipped the observation so the high reward is at observation 0 and the lower reward (+0.005 versus +1 upstream) at observation 1, to remove an inadvertent bias from standard random initialization and ReLU activation that it says made BDQN look artificially good. In Deepsea, taking the action to go right costs 0.01/N except in the bottom-right corner where it yields a reward of 1, and a uniformly random policy reaches the goal with probability 2^{-N} per episode.

Why This Matters

Impact on research. The paper targets a specific and practical gap: directed exploration methods exist and often perform well, but simple undirected ε-greedy remains widely used because directed methods are hard to bolt onto existing pipelines. VBE aims to make optimism cheap to adopt — one extra network update per step, no change to the base learning algorithm. It also reframes the relationship between random-prior ensembles (BDQN) and reward-bonus ensembles (RND/ACB), showing that random-prior style bonuses can be reinterpreted as values of a random reward, which may open new ways to reason about and design optimism.

Real-world applications. The paper does not name specific application domains; these are generic settings where RL exploration is used and where the paper's properties would matter:

  • Robotics and control, where the paper's classic-control suite (Mountain Car, Puddle World) is a standard proxy and where sparse rewards make exploration the bottleneck.
  • Game playing and simulation-based decision making, which the Atari experiments directly represent.
  • Sequential decision problems with expensive data collection, where sample efficiency from directed exploration directly reduces cost.
  • Any domain with irreversible or costly mistakes — such as clinical dosing or process control — where an easy-to-tune, drop-in exploration bonus is preferable to a full algorithm replacement.

Industry relevance. The pitch is deliberately practical: compatibility with any base RL algorithm, a small additional computational footprint per step, and only two parameters discussed in the sensitivity analysis (ensemble size and bonus scale). That lowers the barrier for teams that currently default to ε-greedy because more sophisticated options are seen as too onerous.

Future Directions

  • Convergence theory under changing policies. The paper explicitly states that in VBE both the behavior policy and the target policy change with time, and that TD theory does not address this scenario; existing results cover fixed behavior policies (Zhao et al., 2021) or DQN variants with a fixed dataset (Wang and Ueda, 2022). The changing behavior policy alters the relative importance of states in the objective, and the paper knows of no theory that would guarantee the bonuses converge to zero.
  • Guaranteeing bonus decay in practice. Although the paper reports that bonuses always converged to zero in its own experiments, establishing conditions under which this is guaranteed remains open — in particular whether off-policy gradient TD methods provably reach global solutions for the network classes of interest.
  • Scaling and deeper evaluation. The Atari results are described as "demonstrative experiments"; broader and harder exploration suites, and further analysis of ensemble size and bonus scale (which the paper places in Appendix E), are natural extensions.
  • Extension beyond Double DQN. The paper notes that the algorithm could be swapped into other off-policy value-based algorithms, and that actor-critic methods with an explicit critic could incorporate the bonuses through an optimistic critic, but restricts its experiments to Double DQN. The choice of target policy for the ensemble updates is investigated separately in Appendix F.

Target Audience

Reinforcement learning researchers and practitioners who work on exploration, sample efficiency, or uncertainty estimation in deep RL. It is most useful to readers already comfortable with DQN-family algorithms, TD bootstrapping, and ensembles, but the central mechanism — train predictors against fixed random value functions and use the residual as an optimism signal — is accessible to graduate students and engineers who want a drop-in alternative to ε-greedy. Readers looking only for headline benchmark numbers will find the truncated content lacking the specific per-environment scores, which are reported in the appendices and figures of the full paper rather than in the text available here.

Authors’ abstract

Optimistic value estimates provide one mechanism for directed exploration in reinforcement learning (RL). The agent acts greedily with respect to an estimate of the value plus what can be seen as a value bonus. The value bonus can be learned by estimating a value function on reward bonuses, propagating local uncertainties around rewards. However, this approach only increases the value bonus for an action retroactively, after seeing a higher reward bonus from that state and action. Such an approach does not encourage the agent to visit a state and action for the first time. In this work, we introduce an algorithm for exploration called Value Bonuses with Ensemble errors (VBE), that maintains an ensemble of random action-value functions (RQFs). VBE uses the errors in the estimation of these RQFs to design value bonuses that provide first-visit optimism and deep exploration. The key idea is to design the rewards for these RQFs in such a way that the value bonus can decrease to zero. We show that VBE outperforms Bootstrap DQN and two reward bonus approaches (RND and ACB) on several classic environments used to test exploration and provide demonstrative experiments that it can scale easily to more complex environments like Atari.

Read the original paper