Research
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
Actor-Free Continuous Control via Structurally Maximizable Q-Functions Overview Research area: Off-policy reinforcement learning for continuous control (value-based Q-learning, actor-free methods). Te

- arXiv
- 2510.18828
- Published
- 2025-10-21
- Authors
- Yigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem Bıyık
AI summary
Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsOverview
Research area: Off-policy reinforcement learning for continuous control (value-based Q-learning, actor-free methods).
Technical level: Advanced. The paper assumes familiarity with Markov Decision Processes, the Bellman equation, DQN-style value iteration, and actor-critic algorithms such as TD3, DDPG, and SAC.
Scope: The paper proposes Q3C, an actor-free Q-learning algorithm for continuous action spaces that learns a Q-function built from a set of learned "control-points" so that the maximizing action can be read off directly, without training a separate policy network.
What This Paper Is About
Value-based RL methods like DQN are simple and stable, but they only work in discrete action spaces because they require maximizing over all actions. Continuous control therefore usually falls back on actor-critic methods (DDPG, TD3, SAC), which train a separate actor network to maximize the critic's output and are often unstable and prone to settling on locally optimal actions. The paper's goal is to build a purely value-based, actor-free algorithm for continuous control by using a Q-function representation whose maximum is structurally guaranteed to sit at one of a small set of learned control-points.
Key Contributions
-
A structurally maximizable Q-function combined with deep learning. The authors revisit the "wire-fitting" control-point interpolation framework (Baird and Klopf) and pair it with deep networks so that the argmax over a continuous action space reduces to an argmax over scalar values, one per control-point.
-
An improved architecture that separates control-point generation from value estimation. Instead of predicting Q-values and control-point actions independently from a shared trunk, Q3C uses a control-point generator that outputs N control-point actions, plus a separate Q-estimator that evaluates those actions, so Q-values are conditioned on the corresponding actions.
-
Algorithmic improvements for robustness across tasks. These include relevance-based control-point filtering (using only the top k weights), a pairwise separation loss to diversify control-points, and scale-aware normalization of control-point values with exponentially annealed smoothing.
-
Evaluation against actor-free and actor-critic baselines. Q3C is compared against TD3, NAF, RBF-DQN, and vanilla wire-fitting on seven standard Gymnasium tasks and three restricted-action variants, with ten random seeds per algorithm.
Main Findings
-
Comparable to TD3 on standard benchmarks. On Pendulum-v1, Swimmer-v4, Hopper-v4, BipedalWalker-v3, Walker2d-v4, and HalfCheetah-v4, Q3C's final returns are on par with TD3. For example, Q3C reaches 3206.14 ± 407.23 on Hopper-v4 versus TD3's 3113.41 ± 888.17, 3977.39 ± 879.70 on Walker2d-v4 versus TD3's 4770.82 ± 560.16, and 9468.66 ± 949.01 on HalfCheetah-v4 versus TD3's 9984.74 ± 1076.58.
-
One clear shortfall. On Ant-v4, Q3C reaches 3698.41 ± 1314.88, below TD3's 5167.68 ± 673.44. The paper describes this as the exception where Q3C performs suboptimally relative to TD3.
-
Large gains over vanilla wire-fitting. Across all standard benchmarks Q3C substantially outperforms the vanilla wire-fitting baseline, which scores 1987.50 ± 1127.06 on Hopper-v4, 2462.30 ± 1095.41 on Walker2d-v4, and 7546.23 ± 1234.31 on HalfCheetah-v4.
-
Q3C outperforms TD3 decisively under restricted action spaces. In environments where the action space is restricted to a set of hyperspheres (actions outside are invalid and have no effect), Q3C achieves 1000 ± 0 on InvertedPendulumBox (TD3: 782.76 ± 348.92), 4357.82 ± 1503.33 on HalfCheetahBox (TD3: 2276.70 ± 2036.59), and 1974.28 ± 1170.05 on HopperBox (TD3: 1406.83 ± 1162.72). The authors attribute TD3's degradation to gradient ascent getting stuck in non-convex Q-landscapes.
-
One restricted-environment exception. The paper states that NAF performs marginally better than Q3C on HalfCheetahBox-v4 (NAF: 4867.05 ± 1487.69 versus Q3C: 4357.82 ± 1503.33), which the authors suggest may indicate that a quadratic Q-function approximation is sufficient there.
-
Other value-based baselines underperform. NAF consistently underperforms, which the paper attributes to its quadratic-in-action inductive bias. RBF-DQN also achieves suboptimal results; the paper notes it requires roughly 100 centroids to reach sufficient expressivity, limiting scalability to high-dimensional action spaces.
-
Every ablation component matters. Removing conditional Q-value generation, ranking/filtering, diversification, or normalization reduces final performance on Hopper-v4, BipedalWalker-v3, Walker2d-v4, and HalfCheetah-v4. The largest degradation comes from removing the diversification loss, which drops BipedalWalker-v3 to -67.8 ± 119.6 and HalfCheetah-v4 to 5282.7 ± 1116.2. Even the weakest ablated variant still outperforms vanilla wire-fitting on these tasks.
-
Universal approximation is preserved. The appendix proves that for any continuous Q-function on a compact action set and any epsilon > 0, a finite set of control-points exists such that the wire-fitting interpolator is within epsilon of the true function in the sup-norm.
-
Evaluation protocol. Each algorithm is trained with 10 random seeds, evaluated every 10,000 steps with 10 rollout episodes; curves report the mean with shading at one standard error across 10 trials.
-
Not reported in the provided content: the specific number of control-points N and the filtering value k, wall-clock training time, and total compute cost.
Methodology in Plain English
The core idea is to change what the Q-network outputs so that finding the best action becomes trivial.
Control-points instead of a smooth surface. The Q-function for a given state is represented by a small set of N "control-point" actions and a corresponding set of N Q-values. The Q-value of any other action is computed by an inverse-distance weighted interpolation between these points. Because the weights are constructed to depend on how much worse each point is than the best point, the maximum of the interpolated function is guaranteed to occur at one of the control-points. So to find the best action, the algorithm only has to compare N scalar Q-values — no gradient ascent, no actor.
Splitting the network. Rather than having one network predict both actions and values, Q3C uses two parts: a control-point generator that proposes N candidate actions, and a Q-estimator that scores each of those actions. This keeps Q-values consistent across control-points that land near each other.
Filtering to the most relevant points. When evaluating Q(s, a) for a given action, only the top k control-point weights are used, discarding the other N-k. This sharpens the local landscape instead of letting distant points blur it.
Keeping control-points spread out. A separation loss encourages control-points to be uniformly dispersed rather than collapsing toward the corners or boundaries of the action space.
Taming scale differences. Action spaces are normalized to [-1, 1], the control-point generator uses a tanh nonlinearity, and each state's control-point values are rescaled to [0, 1] inside the weight term. The smoothing factor is annealed exponentially so large rewards do not drown out spatial information.
Borrowing TD3's stabilization toolkit. Q3C uses Gaussian exploration noise, twin Q-networks to reduce overestimation, target networks for stationary targets, and target policy smoothing on the maximizing action. Training alternates between acting in the environment, storing transitions, sampling minibatches, computing Bellman and separation losses, and periodically updating target networks.
Why This Matters
The work shows that the long-standing assumption that continuous control requires an actor can be relaxed: a Q-network can be built so that maximization over the action space is cheap and exact-by-construction. This reopens a direction (deep wire-fitting) that had been largely abandoned due to poor benchmark performance, and it matters most precisely where actor-critic methods break down — action spaces with discontinuities or safety constraints.
Real-world applications:
- Safe robotics and constrained actuation. Tasks where only a subset of commands is admissible (joint limits, safety envelopes) create non-smooth value landscapes that defeat gradient-based actors.
- Industrial process control. Systems with restricted, discrete-ish or bounded actuator ranges where locally optimal policies are costly.
- Autonomous vehicles and drones. Control under hard physical or regulatory action constraints, where non-convex value functions are common.
- Simulation-based control pipelines (e.g., locomotion). The benchmark tasks — HalfCheetah, Hopper, Walker2d, Ant — map directly onto legged locomotion, a domain where sample-efficiency failures and training instability are practical bottlenecks.
Industry relevance: the method is actor-free, meaning it removes a whole network, its hyperparameters, and its inference-time cost, which lowers tuning burden and deployment latency. It also drops into an existing TD3-style training stack, as the authors demonstrate by implementing Q3C on the stable-baselines3 backbone.
Future Directions
- Better exploration. Q3C simply inherits TD3's exploration scheme, and the paper notes its sample efficiency can lag behind other baselines in some environments. The authors suggest options such as a Boltzmann distribution over control-point values.
- Borrowing sample-efficiency tricks. n-step returns, prioritized experience replay, or batch normalization layers in the critic instead of target networks are named as promising additions.
- Offline RL. The authors suggest that the interpolation constraints inherent to control-points could naturally mitigate Q-value overestimation, a known difficulty in offline learning.
- Stochastic policies. The current work uses Q3C only in deterministic Q-learning; extending the control-point architecture to model a soft Q-function, in the style of SAC, is left open.
Target Audience
Reinforcement learning researchers and graduate students working on continuous control, value-based RL, or actor-critic stability; practitioners who need constrained-action control without the tuning burden of an actor network; and readers interested in function-approximation design, since the paper's central argument is about how to shape a Q-network so that maximization becomes structural rather than numerical.
Authors’ abstract
Value-based algorithms are a cornerstone of off-policy reinforcement learning due to their simplicity and training stability. However, their use has traditionally been restricted to discrete action spaces, as they rely on estimating Q-values for individual state-action pairs. In continuous action spaces, evaluating the Q-value over the entire action space becomes computationally infeasible. To address this, actor-critic methods are typically employed, where a critic is trained on off-policy data to estimate Q-values, and an actor is trained to maximize the critic's output. Despite their popularity, these methods often suffer from instability during training. In this work, we propose a purely value-based framework for continuous control that revisits structural maximization of Q-functions, introducing a set of key architectural and algorithmic choices to enable efficient and stable learning. We evaluate the proposed actor-free Q-learning approach on a range of standard simulation tasks, demonstrating performance and sample efficiency on par with state-of-the-art baselines, without the cost of learning a separate actor. Particularly, in environments with constrained action spaces, where the value functions are typically non-smooth, our method with structural maximization outperforms traditional actor-critic methods with gradient-based maximization. We have released our code at https://github.com/USC-Lira/Q3C.