Research
Higher-Order Action Regularization in Deep Reinforcement Learning: From Continuous Control to Building Energy Management
Overview Research area: Deep reinforcement learning (RL) for continuous control, with a focus on action smoothness regularization and its transfer to building energy management (HVAC control). The pap

- arXiv
- 2601.02061
- Published
- 2026-01-05
- Authors
- Faizan Ahmed, Aniket Dixit, James Brusey
AI summary
Overview
Research area: Deep reinforcement learning (RL) for continuous control, with a focus on action smoothness regularization and its transfer to building energy management (HVAC control). The paper was presented in the context of the UrbanAI workshop on "Harnessing Artificial Intelligence for Smart Cities," and comes from the Centre for Computational Science and Mathematical Modelling at Coventry University.
Technical level: Intermediate. Readers need some familiarity with Markov decision processes, policy gradient methods such as PPO, and the notion of numerical derivatives, but the central idea (penalizing jerky actions) is intuitive.
Scope: The paper systematically compares first-, second-, and third-order action derivative penalties in deep RL, first on four continuous control benchmarks and then on a two-zone HVAC control environment.
What This Paper Is About
Standard reinforcement learning maximizes cumulative reward without regard for how abruptly an agent changes its actions, which produces erratic, high-frequency control that wastes energy and stresses physical equipment. This paper asks whether adding penalties on the first, second, or third derivative of the action sequence to the reward function can produce smoother policies, and which derivative order gives the best smoothness-versus-performance trade-off. It then tests that question beyond simulated benchmarks, in a building HVAC control setting where smooth actions translate directly into less equipment switching.
Key Contributions
- A systematic derivative-order comparison. The authors introduce a third-order (jerk) penalty and evaluate it alongside first-order (velocity) and second-order (acceleration) penalties on the same continuous control tasks, using identical penalty weights (λ₁ = λ₂ = λ₃ = 0.1) so that derivative order is not confounded with scale.
- Evidence that jerk minimization is the best trade-off. Across the four benchmarks, third-order penalties consistently produce the smoothest policies while retaining competitive task performance.
- An action-history state augmentation. The state is extended to include the three previous actions alongside the current state, which preserves the Markov property while making derivative computation possible directly from the policy's input.
- Transfer to building energy management. The method is validated on a custom, open-source "DollHouse" two-zone HVAC environment with SINDy-identified dynamics, showing a 60% reduction in equipment switching events.
Main Findings
- Smoothness improves monotonically with derivative order. First-order penalties give modest smoothness gains, second-order penalties strengthen the effect, and third-order penalties give the largest reductions in action jerkiness in every environment tested.
- Jerk standard deviation dropped sharply under third-order penalties relative to the baseline: 78.8% in HalfCheetah, 77.3% in Hopper, 38.6% in Reacher, and 58.3% in LunarLander.
- Third-order HalfCheetah: reward 725 ± 70 with smoothness 1.443 ± 0.030, versus baseline reward 1052 ± 146 with smoothness 6.806 ± 0.098.
- Third-order Hopper: reward 1379 ± 845 with smoothness 1.822 ± 0.077, versus baseline reward 1977 ± 629 with smoothness 8.012 ± 0.150.
- Third-order Reacher: reward −5 ± 2 with smoothness 0.097 ± 0.010, versus baseline reward −6 ± 2 with smoothness 0.158 ± 0.017.
- Third-order LunarLander: reward 230 ± 68 with smoothness 1.053 ± 0.091, versus baseline reward 204 ± 90 with smoothness 2.524 ± 0.122 — the one environment where the third-order variant also improved mean reward.
- Smoothness costs performance in some environments. The authors characterize the reward loss as modest and note that it is not uniform across tasks; in Reacher, for example, intermediate-order results were close to the baseline (first-order reward 102 ± 125 and smoothness 2.224 ± 0.122).
- HVAC switching fell by 60%. Smooth control reduced equipment switching events, which the paper links to reduced startup energy penalties, better thermal efficiency, and longer equipment life.
- Contextual energy figures. The paper cites prior work reporting that on-off control increases energy costs compared to continuous operation, and that demand-limiting control methods achieve energy reduction ratios of 9.8% to 10.5%.
- Three proposed explanations for why higher-order penalties work. Physical alignment (jerk relates to mechanical and thermal stress), learning stability (reduced gradient noise during training), and practical deployment (better transfer when actuator dynamics and sensor noise are present).
- What is not reported. The paper does not quantify monetary savings, measured equipment lifetime extension, or comfort violations in the HVAC environment; those effects are described qualitatively.
Methodology in Plain English
The authors frame continuous control as a Markov decision process, then make one modification: they let the policy see not just the current state but also the last three actions it took. That way, the "jerkiness" of the action sequence — how much the rate of change of the action is itself changing — can be computed and penalized directly inside the reward.
They define three penalty terms, one per derivative order, using finite differences: a first-order term on the difference between consecutive actions, a second-order term on the acceleration-like combination, and a third-order term on the jerk-like combination. Each penalty is subtracted from the reward, scaled by a weight λ. All three weights are set to the same value, 0.1, across every experiment so that the comparison isolates the effect of derivative order rather than the effect of penalty strength.
They then train PPO agents for 1 million timesteps on four OpenAI Gym environments — HalfCheetah-v4, Hopper-v4, Reacher-v4, and LunarLanderContinuous-v2 — with the same hyperparameters for every method and results averaged over 5 random seeds. Smoothness is measured as the standard deviation of jerk computed via third-order finite differences. Finally, they apply the same regularization idea to a custom two-zone HVAC control environment, DollHouse, whose dynamics were identified from building data using SINDy, controlling temperature setpoints and damper positions to balance energy use against occupant comfort.
Why This Matters
Impact on research. The paper reframes smoothness not as a reward-engineering afterthought but as a derivative-order design choice with a measurable ordering of outcomes. It supplies a common protocol — matched penalty weights, matched hyperparameters, five seeds — for comparing regularization orders, and it connects RL practice to older robotics and control-theory ideas about jerk, total variation, and switching losses. It also highlights a gap: most RL work on building control has focused on reward shaping rather than fundamental smoothness constraints.
Real-world applications:
- Building HVAC control: reducing compressor and fan cycling, which the paper argues lowers startup energy penalties and extends the finite switching cycles of components.
- Robotic manipulators: jerk minimization addresses mechanical stress and vibration from abrupt movements, a problem noted in prior robotics work the paper cites.
- Spacecraft and fuel-constrained control: the LunarLander environment represents fuel-efficient vehicle control, where smooth thrust changes matter.
- Unstable balancing and precise manipulation systems: Hopper and Reacher stand in for platforms where large action swings degrade stability and precision.
Industry relevance. As building automation and smart-city infrastructure increasingly adopt learned controllers, operational constraints such as actuator wear, startup energy costs, and equipment lifetime become deployment blockers rather than simulation details. A regularization term that is simple to implement and reduces switching by 60% in a two-zone testbed is directly relevant to building management system vendors, energy service companies, and facility operators evaluating RL-based control. The paper's framing of smoothness as a deployment precondition rather than a tuning preference is the distinctive angle.
Future Directions
- Adaptive penalty weighting. The authors explicitly call for schemes that adjust smoothness constraints based on system state and operational requirements, replacing the fixed λ used here.
- Principled hyperparameter selection. Penalty magnitudes currently require domain-specific tuning; systematic methods for choosing them across applications remain an open challenge, as does deciding when smoothness should be prioritized over raw performance.
- Integration with model-predictive control. Combining these penalties with MPC for longer planning horizons is proposed.
- Broader energy-critical scope. Validation is limited to HVAC; the authors identify lighting, elevators, and other building systems as untested territory.
Target Audience
RL researchers working on safe or deployable control, especially those interested in regularization and constrained policy optimization; control engineers and building energy specialists evaluating learned controllers for HVAC and other physical systems; and practitioners in smart-city, building automation, or robotics settings who need policies that respect actuator and energy constraints. It is most useful to readers who already understand policy gradient methods and want a concrete, comparative result about which smoothness penalty to use and what it costs in task performance.
Authors’ abstract
Deep reinforcement learning agents often exhibit erratic, high-frequency control behaviors that hinder real-world deployment due to excessive energy consumption and mechanical wear. We systematically investigate action smoothness regularization through higher-order derivative penalties, progressing from theoretical understanding in continuous control benchmarks to practical validation in building energy management. Our comprehensive evaluation across four continuous control environments demonstrates that third-order derivative penalties (jerk minimization) consistently achieve superior smoothness while maintaining competitive performance. We extend these findings to HVAC control systems where smooth policies reduce equipment switching by 60%, translating to significant operational benefits. Our work establishes higher-order action regularization as an effective bridge between RL optimization and operational constraints in energy-critical applications.