Research
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Overview Research area: Reinforcement learning — specifically exploration control and entropy regularization in non-stationary (drifting) environments. Technical level: Advanced. The paper is built ar

- arXiv
- 2601.19624
- Published
- 2026-01-27
- Authors
- Tongxi Wang, Zhuoyang Xia, Xinran Chen, Shan Liu
AI summary
Overview
- Research area: Reinforcement learning — specifically exploration control and entropy regularization in non-stationary (drifting) environments.
- Technical level: Advanced. The paper is built around online convex optimization, dynamic regret, mirror descent, and maximum-entropy RL theory, with deep RL experiments layered on top.
- Scope: The paper proposes AES (Adaptive Entropy Scheduling), a plug-in mechanism that scales the entropy weight/temperature online using an observable drift proxy, derives a square-root scaling rule for that weight from a per-round tracking-versus-stability trade-off, and evaluates it on four algorithm carriers across three task families under injected environment drift.
What This Paper Is About
Most maximum-entropy RL methods treat the entropy coefficient or target entropy as fixed, which causes over-exploration during stable periods and under-exploration after the environment changes. This paper asks how exploration intensity should scale with the magnitude of environmental change, and answers it by casting entropy scheduling as a dynamic-regret trade-off between tracking a drifting optimal policy and avoiding unnecessary randomness. It then turns that principle into an online rule driven by a training signal already available for free (the upper quantile of absolute TD errors).
Key Contributions
-
A principled formulation of entropy scheduling. Under standard assumptions, entropy scheduling in non-stationary maximum-entropy RL is cast as the dynamic-regret trade-off between tracking a drifting comparator and stabilizing updates, reducing the design question to a single control variable, λ_t (the entropy strength).
-
A square-root scaling rule. Theorem 3.2 shows the per-round loss takes the form C₁ ξ_t/λ_t + C₂ λ_t, where ξ_t is comparator drift. Minimizing gives the oracle scale λ_t* = √(C₁/C₂ · ξ_t), i.e., entropy weight grows like the square root of drift.
-
A fully online schedule from an observable proxy. Theorem 3.5 replaces the unobservable drift ξ_t with a conservative proxy ξ̂_t satisfying ξ_t ≤ ξ̂_t, giving λ_t = √(C₁/C₂) √(Â_t/t), clipped to [λ_min, λ_max] in practice, with regret bounded by 4√(C₁C₂ T Â_T) plus a constant.
-
A bridge to non-stationary soft RL and a plug-in implementation. Theorem 3.6 links MDP variation to drift of the soft-optimal comparator and gives a planning-version principal-term bound. AES is then integrated into SAC, PPO, SQL, and MEow without changing the underlying actor–critic or value-learning structure, replacing only the fixed entropy weight.
Main Findings
-
Drift is measurable for free. The default drift proxy is the q-quantile of absolute TD errors with q = 0.9, smoothed by a window or EMA. The paper states this is an inherent training parameter obtainable without extra cost or operation, and that AES incurs negligible overhead (one scalar statistic plus a lightweight scheduler update per iteration).
-
Oracle entropy scale is √ξ_t. Theorem 3.3 gives λ_t* = √(C₁/C₂ · ξ_t), and the resulting regret scales as 2√(C₁C₂) √(T ∑ξ_t) plus a constant, i.e., λ_t* ≍ √ξ_t at dominant order.
-
MDP drift propagates to policy drift. The non-stationarity budget is B_T^MDP = ∑(Δ_t^r + γ V_max Δ_t^P), with ‖Q_t* − Q_{t−1}*‖_∞ ≤ (1/(1−γ))(Δ_t^r + γ V_max Δ_t^P) and policy drift bounded by (1/μ) times the Q drift. The per-state soft-RL gap equals μ times a KL divergence.
-
Consistent empirical gains across carriers and families. In Table 2 (normalized AUC, higher is better), SAC under Mixed drift averaged 0.73 on Toy, 0.65 on MuJoCo, and 0.51 on Isaac Gym, while SAC+AES reached 0.97, 0.94, and 0.79 respectively. On Isaac Gym under Periodic drift, SAC scored 0.57 and SAC+AES scored 0.95. MEow+AES reached 1.03 on Toy Abrupt and 1.02 on Toy Linear.
-
Drop-area ratio generally falls. For example, MEow+AES reported 0.03 on MuJoCo Abrupt and 0.09 on Isaac Gym Periodic. Negative values are possible when a non-stationary run outperforms its steady reference — PPO+AES on Toy Abrupt (−0.03) and Linear (−0.03), and MEow+AES on Toy Abrupt (−0.06), Linear (−0.05), and Periodic (−0.04).
-
Steady performance is not systematically harmed. Table 2 shows normalized AUC under Steady that is comparable or slightly improved for AES variants, which the authors read as evidence that AES is not a change-event-specific patch but a general calibration of exploration strength.
-
Faster recovery after abrupt changes. In Table 3 (fraction of total training steps needed to return to the pre-change level), Toy 2d multi-goal dropped from 7.8% for SAC to 3.7% for SAC+AES and from 6.1% for MEow to 3.2% for MEow+AES. MuJoCo Humanoid went from 14.8% (SAC) to 8.6% (SAC+AES). Isaac Gym AllegroHand went from 21.5% (SAC) to 10.6% (SAC+AES), and FrankaCabinet from 19.4% (SAC) to 10.5% (SAC+AES).
-
The RL result is deliberately scoped. Theorem 3.6 is described as a planning-version principal-term result, not a complete finite-sample theorem for deep off-policy learning; approximation, sampling, and occupancy effects are kept as explicit Bias_t residuals.
-
Not reported in the available content. The number of random seeds, dispersion values, wall-clock or compute cost figures, and any results beyond the start of Section 4.4 are not present in the provided text. The paper is truncated after Table 3.
Methodology in Plain English
The argument proceeds in three steps.
First, the authors deliberately strip away RL complexity and work in a simpler setting: online convex optimization over a probability simplex, where the target solution moves over time. They measure how much the target moves round to round (the drift ξ_t) and use entropy mirror descent to update the decision. A dynamic mirror descent inequality isolates drift as a sum of drift terms divided by the learning rate.
Second, they add entropy regularization to the loss and then couple the learning rate to the entropy weight (η_t = c λ_t). After substituting bounds and absorbing higher-order terms into constants, the per-round loss separates into two competing pieces: a tracking cost that grows when λ is too small, and a stability cost that grows when λ is too large. Minimizing that expression over λ gives the square-root rule. Since true drift is unobservable, they substitute a conservative observable proxy and take a running average, producing a schedule that is a square root of the average proxy value so far, clipped to a fixed range.
Third, they connect this to RL: environmental changes in rewards and transitions perturb the soft-optimal Q-function, which in turn perturbs the soft-optimal policy, so comparator drift in the abstract setup corresponds to real MDP drift. This justifies using the upper quantile of TD errors — which rises when previously accurate value predictions go stale — as the proxy. No change-point detection or restart is used; entropy is modulated continuously.
For experiments, they inject controlled drift (goal changes in the 2D multi-goal Toy domain, dynamics or target variations in MuJoCo, and scalable physics or task perturbations in Isaac Gym) across Steady, Abrupt, Linear, Periodic, and Mixed patterns. Note that the abstract counts "4 drift modes" while the experimental protocol lists five patterns including the no-drift Steady reference. All methods are aligned by training progress rather than wall-clock time so that every method meets the same drift events at the same fraction of environment steps. Three metrics are reported: normalized AUC (AUC divided by the Steady SAC baseline within the same task family), performance drop area ratio (1 minus the non-stationary AUC over the steady AUC), and abrupt-change recovery time normalized by total training steps. The SAC baseline used automatic temperature control.
Why This Matters
Impact on research. Prior variation-aware analyses were largely decoupled from modern maximum-entropy deep RL, and adaptive entropy methods in practice were heuristic or tuned per environment. This work supplies an explicit, closed-form link between non-stationarity and the exploration knob, and it separates the dominant non-stationarity term from interface error terms — a decomposition that helps identify what actually needs to be controlled in a training pipeline.
Real-world applications (as described or implied by the paper's task choices):
- Robotics, where physical conditions change over time and agents must re-adapt rather than commit to an outdated policy.
- Autonomous driving, which must cope with evolving traffic patterns.
- Recommender systems, which must continuously adapt to shifting user preferences.
- Large-scale simulated locomotion and manipulation (legged robots such as Ant, Humanoid, ANYmal, and dexterous hands such as AllegroHand and FrankaCabinet), where randomized physics makes drift a routine rather than exceptional condition.
Industry relevance. AES requires almost no structural changes, adds only a scalar TD-error statistic and a scheduler update, covers both off-policy (SAC, SQL, MEow) and on-policy (PPO) learning, and needs no change-point detector or restart machinery. That combination — low integration cost with a defensible theoretical rationale — is what makes it practically deployable where heuristic entropy tuning currently dominates.
Future Directions
- From planning to full learning guarantees. Theorem 3.6 is explicitly a planning-version principal-term bound; extending it to a complete finite-sample result for deep off-policy learning with function approximation, sampling, and occupancy residuals would be the natural next theoretical step.
- Better calibrated and alternative drift proxies. The paper notes that model disagreement or critic-parameter drift would also fit the same interface, and that the proxy only needs to be conservative, not unbiased. Calibrating and validating alternative signals is an open direction.
- Combining with interface-stability methods. The related-work discussion explicitly suggests pairing AES with techniques that stabilize the critic, replay, or optimizer (for example clustered replay and corrective gradient penalties, and analyses of TD feedback dynamics) so that interface errors are reduced while AES adapts exploration.
- Broader evaluation of the adaptation claim. The authors frame the experiments as testing the structural prediction that exploration should rise with non-stationarity, not as numerical verification of the regret bounds. Extending the study to more drift intensity levels, additional proxies, and further domains (for example multi-agent or language-model RL, where adaptive entropy has recently been explored) would probe whether the square-root scaling holds in practice.
Target Audience
This paper is best suited to reinforcement learning researchers and graduate students working on non-stationary or continual RL, exploration and entropy regularization, and the theory of dynamic regret. It is also relevant to practitioners who maintain deployed RL systems — robotics, autonomy, recommendation and control engineering — and who need a low-overhead, theoretically motivated alternative to repeated manual tuning of entropy coefficients or target entropy. Readers without a background in online convex optimization or maximum-entropy RL will find the theory sections demanding, though the experimental section and Table 1's integration guide are accessible on their own.
Authors’ abstract
Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift, and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude. We show that, under standard assumptions, entropy scheduling in non-stationary maximum-entropy RL can be cast as the dynamic-regret trade-off between tracking a drifting comparator and stabilizing updates, yielding a square-root scaling rule for the entropy weight in terms of a online non-stationarity proxy. Building on this, we propose AES--Adaptive Entropy Scheduling--which adaptively adjusts the entropy coefficient/temperature online using observable drift proxies during training, requiring almost no structural changes and incurring minimal overhead. Across 4 algorithm variants, 12 tasks, and 4 drift modes, AES significantly reduces the fraction of performance degradation caused by drift and accelerates recovery after abrupt changes.