Research
Online Action-Stacking Improves Reinforcement Learning Performance for Air Traffic Control
Overview Research area: Reinforcement learning (RL) applied to air traffic control (ATC), specifically the design of action spaces and inference-time techniques that make RL policies operationally rea

- arXiv
- 2601.04287
- Published
- 2026-01-07
- Authors
- Ben Carvell, George De Ath, Eseoghene Benjamin, Richard Everson
AI summary
Overview
Research area: Reinforcement learning (RL) applied to air traffic control (ATC), specifically the design of action spaces and inference-time techniques that make RL policies operationally realistic.
Technical level: Intermediate. The paper assumes familiarity with Markov Decision Processes, policy and value functions, actor-critic methods, and Proximal Policy Optimisation (PPO), though the core idea (grouping small commands into one larger command) is explained in accessible terms.
Scope: The paper introduces and evaluates "online action-stacking", an inference-time wrapper that compiles bursts of primitive 10-degree heading or 10-flight-level actions into compound ATC-style clearances, trained with PPO on the BluebirdDT digital twin across lateral navigation, vertical navigation, and two-aircraft collision avoidance scenarios.
What This Paper Is About
RL agents for ATC typically control aircraft through small incremental adjustments, so achieving a manoeuvre that a controller would issue as one instruction ("turn left 30 degrees") requires many consecutive commands ("turn left 10 degrees", three times). Real ATC practice demands the fewest clearances necessary, to manage pilot workload, controller workload, and congestion on the radio frequency. The paper's goal is to let agents train on a very small discrete action space while still producing domain-appropriate compound clearances, by stacking repeated primitive actions at inference time.
Key Contributions
- Online action-stacking: an inference-time procedure that repeatedly queries a trained policy for the same aircraft, accumulates identical 10-degree increments (left or right) into a single macro-command, and stops when the policy outputs "no action" or acts on the other aircraft. It leaves the training MDP unchanged and performs multiple policy queries at a single physical time step, holding the environment state fixed so the Markov structure is preserved.
- An action-damping reward that induces bursts: the damping penalty decays slowly after an action (with a threshold of n_max = 10 steps, equivalent to one minute between clearances), which incentivises issuing consecutive commands with no gaps and produces the action bursts that stacking can compile.
- A demonstration that a 5-dimensional action space can match a 37-dimensional one: for two-aircraft lateral navigation, the damped-and-stacked policy achieves comparable performance to a policy trained with a 37-action space, despite operating with only five actions.
- Domain-aligned scenario suite in BluebirdDT: scenarios for lateral route navigation, vertical navigation to a coordinated exit level, and two-aircraft lateral collision avoidance under the 5 nautical mile separation standard, all run in the "X-Plus Sector" airspace.
Main Findings
- Undamped policies are unusable for stacking: trained for 2 million time steps with only the centreline distance reward plus the terminal reward set, the policy issued a mean of 113.0 actions per episode (standard deviation 15.9) across a sample of 100 episodes, with frequent oscillations between left and right commands. The example shown in the paper issued 92 separate actions.
- Action damping cuts instruction counts substantially: adding the action damping penalty produced a mean of 14.5 actions per episode (standard deviation 6.6) over 100 episodes, an 87% reduction compared with the undamped policy. The example figure shows 15 separate actions.
- Action stacking roughly halves the remaining commands: the damped and stacked policy produced a mean of 7.2 actions per episode (standard deviation 3.5) across 100 episodes, a mean reduction of approximately 50%. The example figure shows 7 separate actions, and a central turn near the fix EGL that the damped policy performed as seven repeated "turn right 10 degrees" actions became a single "turn right 70 degrees" clearance.
- Comparable performance with a far smaller action space: the policy trained with the large 37-action space issued a mean of 6.3 actions per episode (standard deviation 3.0) over 100 episodes. Although this is lower than the 7.2 of the stacked policy, the paper reports the difference as not statistically significant (standard deviations of approximately 3). The 5-dimensional action space performed similarly to the 37-dimensional one.
- Faster convergence with the smaller action space: reward convergence was faster for the damped lateral navigation policy than for the large action space policy, though both trained to a reasonable level of convergence at approximately the same reward level on this simpler problem.
- The instruction-count criterion drove failures for the large action space policy: the success rate (satisfying all three terminal criteria: navigation to exit, no airspace excursions, and fewer than 30 actions issued) became reliable much later for the large action space policy, and inspection showed the number-of-actions criterion caused the drop, despite the theoretically smaller number of actions needed to complete the scenario.
- Vertical control was investigated but results are not reported in the available content: a simple policy was trained for 2 million time steps to navigate aircraft to a target flight level, and the text introduces figures showing the effect of applying action stacking in the vertical case, but the paper content provided is truncated at that point, so the vertical results are not reported here.
Methodology in Plain English
The researchers built ATC-style scenarios in BluebirdDT, the digital twin platform developed as part of Project Bluebird, using a simplified implementation of the Eurocontrol BADA aircraft performance model. The simulator updates aircraft states every six seconds, matching operational radar update rates in the UK, and each scenario runs for 300 steps (30 minutes of simulated time).
Three scenario families were used:
- Lateral navigation with two Boeing 757-300 aircraft, both at Flight Level 300 (30,000 feet), starting at random positions within a 4 nautical mile radius circle centred on their first route fix, following a randomly chosen viable route through the X-Plus Sector.
- Lateral navigation and avoidance with two aircraft, adding relative bearing and separation distance to the state and a safety reward penalising projected losses of the 5 nautical mile separation standard.
- Vertical navigation with a single aircraft, with initial and target levels each randomly chosen between Flight Level 100 and Flight Level 300.
The agent is a centralised (not multi-agent) PPO policy: a fully connected neural network with 2 hidden layers of 64 neurons each and ReLU activations, learning rate 0.0001, discount factor 0.99, and entropy coefficient 0.01. Training ran for 2 million time steps; avoidance policies were pre-trained for 2 million steps without the safety reward and then trained for a further 30 million steps with it.
For a single aircraft the lateral action space is just three options: no action, turn left 10 degrees, turn right 10 degrees (five actions for two aircraft). Vertical control likewise uses no action, descend 10 flight levels, and climb 10 flight levels. Rewards combine a centreline distance term, an action damping term, a safety term, a vertical term, and a set of terminal bonuses (5 per satisfied criterion, plus 20 when all three are met), with weights configured per scenario.
At inference, action stacking repeatedly asks the policy for another command for the same aircraft, accumulates identical 10-degree increments into one macro-command, and stops when the policy chooses "no action" or switches to the other aircraft. Because all these queries happen at one physical time step with the environment frozen, the approach avoids the partial observability that arises when a macro-action executes over multiple unobserved time steps.
Why This Matters
Impact on research: The paper argues that most RL-for-ATC literature restricts the problem — omitting vertical control, using simplified reward designs, or ignoring the operational need to minimise clearances — which limits near-term relevance. It shows that action sparsity, large action spaces, and domain-accurate instructions can be tackled together by moving the complexity from the action space to an inference-time wrapper, without changing the training objective or introducing semi-Markov formulations.
Real-world applications:
- Decision support for en-route controllers, where an agent could propose safe, orderly, expeditious clearances consistent with CAP493 tactical control requirements.
- Climb and descent management, ensuring aircraft reach the exit level coordinated with the next sector.
- Tactical conflict resolution between aircraft under a minimum separation constraint.
- Reducing radio frequency congestion and pilot and controller workload by issuing fewer, compound clearances rather than streams of small corrections.
Industry relevance: The work sits inside Project Bluebird, an EPSRC Prosperity Partnership between NATS, The Alan Turing Institute, and the University of Exeter, and uses a platform planned for open source release. It responds to pressures on the domain including IATA's estimate of 5.2 billion passengers flying in 2025, rising to as much as 12.4 billion by 2050, alongside UK "Jet Zero" and IATA "Fly Net Zero" commitments to net zero aviation emissions by 2050. It also positions itself relative to existing operational and research systems such as iFACTS, iTEC, ARGOS, SESAR's AGENT and HYPERSOLVER, and the Machine Basic Training framework.
Future Directions
- Scaling to more complex control scenarios. The paper states that action stacking provides a simple mechanism for scaling, and suggests training complexity could be reallocated to higher traffic densities or more robust safety definitions without prohibitive computational cost.
- Completing the vertical control evaluation. The vertical experiments are introduced but the reported content is truncated; the results of stacking for descent and climb, where large flight level changes are commonplace, are not reported here.
- Extending to fuller 3D control. The current formulations treat lateral navigation and avoidance, and vertical navigation separately; combining lateral, vertical, and speed control with wider sector goals remains an open direction relative to the literature the paper reviews.
- Addressing weaker points of the large action space policy. The large action space policy showed inefficient turns and over-correction that the authors suggest could be remedied with further training, and it failed the fewer-than-30-actions success criterion more often — questions remain about how the comparison scales to harder tasks.
Target Audience
RL researchers working on sequential decision-making with action sparsity or large action spaces; ATC researchers and human factors specialists interested in domain-accurate automation; engineers building controller decision-support tools; and industry practitioners in air navigation service providers, research programmes, and regulators who need to judge whether RL policies are operationally realistic rather than only effective.
Authors’ abstract
We introduce online action-stacking, an inference-time wrapper for reinforcement learning policies that produces realistic air traffic control commands while allowing training on a much smaller discrete action space. Policies are trained with simple incremental heading or level adjustments, together with an action-damping penalty that reduces instruction frequency and leads agents to issue commands in short bursts. At inference, online action-stacking compiles these bursts of primitive actions into domain-appropriate compound clearances. Using Proximal Policy Optimisation and the BluebirdDT digital twin platform, we train agents to navigate aircraft along lateral routes, manage climb and descent to target flight levels, and perform two-aircraft collision avoidance under a minimum separation constraint. In our lateral navigation experiments, action stacking greatly reduces the number of issued instructions relative to a damped baseline and achieves comparable performance to a policy trained with a 37-dimensional action space, despite operating with only five actions. These results indicate that online action-stacking helps bridge a key gap between standard reinforcement learning formulations and operational ATC requirements, and provides a simple mechanism for scaling to more complex control scenarios.