Skip to content
AI.info

Research

Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework

Overview Research area: Human-Computer Interaction / AI safety evaluation, applied to Air Traffic Control (ATC) automation. Technical level: Intermediate. The paper is written for a mixed audience of

Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework
arXiv
2601.04288
Published
2026-01-07
Authors
Ben Carvell, Marc Thomas, Andrew Pace, Christopher Dorney, George De Ath, Richard Everson, Nick Pepper, Adam Keane, Samuel Tomlinson, Richard Cannon

AI summary

Overview

Research area: Human-Computer Interaction / AI safety evaluation, applied to Air Traffic Control (ATC) automation.

Technical level: Intermediate. The paper is written for a mixed audience of AI researchers and aviation domain experts; it assumes familiarity with the ATC task but explains the statistical and machine learning methods it references at a high level.

Scope: The paper presents Machine Basic Training (MBT), a human-in-the-loop assessment framework that adapts NATS' regulator-driven "Area Basic" controller training curriculum so that AI agents can be trained and examined against the same competency standards as human trainee controllers.

What This Paper Is About

ATC is a safety-critical, manual decision-making task: Air Traffic Control Officers issue mandatory instructions to aircraft to keep them separated while maintaining efficient flight trajectories. Despite more than five decades of automation research, no system automates this decision-making in operations, and the authors argue a major reason is the misalignment between simplified academic representations of ATC and the real operational environment. The goal is to close that gap by evaluating AI agents through a legally regulated, instructorassessed curriculum rather than purely numerical metrics.

Key Contributions

  1. Machine Basic Training (MBT): A human-in-the-loop training and assessment framework that translates the NATS "Area Basic" course curriculum — the same course taught to new trainee ATCOs — into a form usable for agents, developed through a year-long series of workshops with operational ATCOs and instructors.

  2. Regulatory grounding and explicit scope definition: The framework traces its objectives from UK Regulation 2015/340 and CAP794 through NATS' internal Verification and Cross Reference Index (VCRI), selecting the 39 objectives measurable through practical assessment, and then refines these into a scoped, machine-readable objective set (detailed in the paper's appendix) using granular ICAO competency descriptions.

  3. A validated simulation basis (BluebirdDT) and an inter-rater reliability study: The authors verify that the BluebirdDT digital twin reproduces trajectories flown by real trainee ATCOs on NATS' training college simulators, and demonstrate that instructors score human and machine runs with comparable consistency.

  4. First assessments of two prototype agents: The rules-based agent Hawk and the optimisation-based agent Falcon were tested, with expert assessor gradings reported for January 2025 and a re-test of Hawk in August 2025.

Main Findings

  • Inter-rater reliability is comparable for humans and machines. Across 19 scenarios (each approximately 30 minutes long) independently assessed by at least 7 instructors, the framework produced a mean Spearman's rho of 0.59 and a Kendall's W of 0.64. Permutation tests showed the measured correlations were clearly distinct from randomised score distributions, and there was no appreciable difference in scoring consistency between machine agents and human trainees.

  • The digital twin reproduces college exercises closely. Trajectory data from trainee ATCO runs on NATS' college simulators, logged from November 20th 2024 to June 20th 2025, was replayed through BluebirdDT. Using acceptance thresholds of 5 flight levels vertically and 2.5 nautical miles laterally (half the separation standard), the proportion of aircraft within threshold was 100.0% for Assessment 1 (18 simulations), 97.9% for Assessment 2 (16 simulations), and 92.1% for Assessment 3 (12 simulations). Mean horizontal error ranged from 0.17 to 0.26 NM; mean vertical error ranged from 0.10 to 0.37 flight levels. Under 8% of aircraft exceeded the threshold for all assessments, and these were manually checked.

  • Agent performance tracked expert feedback. In the initial January 2025 round, both agents scored above the minimum in all assessed areas but received mostly unsatisfactory overall gradings. Falcon achieved Satisfactory in Coordination because exit coordination was imposed as a constraint in its problem formulation. Hawk scored more strongly in Controlling due to its expert-informed rule origins.

  • Targeted revisions moved Hawk to Satisfactory in three competencies. After logic revisions and re-testing in August 2025, Hawk was graded Satisfactory in Controlling, Planning, and Coordination, but remained Unsatisfactory in Safety. Falcon's only Satisfactory grade in round 1 was Coordination.

  • Safety remains the unsolved competency. Assessor comments noted that Falcon had "no losses of separation" but "multiple scenarios where safety was not ensured," and Hawk showed one unsafe clearance and six failures to ensure separation. The authors emphasise that the training standard rests on ensuring separation — fail-safe clearances robust to weather, pilot response time, and other uncertainties — not merely predicting conflicts from current trajectories.

Methodology in Plain English

The researchers began with a task that is already regulated: NATS trains new controllers through a 5-month "Area Basic" course using approximately 30 graded simulator exercises, each scored by an instructor against six competency areas — Safety, Controlling, Planning, Coordination, Communication, and Human Factors — plus a set of hidden summative exercises used for final examination.

They rebuilt this course inside BluebirdDT, a configurable probabilistic digital twin of the UK ATC environment that simulates the fictional Medway sector, aircraft performance models, and action spaces. They then worked through the regulatory documents to identify which objectives could realistically be assessed for a machine, and ran 15 workshops with operational controllers and instructors to decide what stays in scope. Safety, Controlling, Planning, and Coordination objectives are assessed; Communication and Human Factors are currently out of scope because radio-telephony is emulated through the simulator API rather than agent speech.

To assess an agent, three 30-minute summative simulations are run with an instructor observing. The instructor writes a report scoring each competency on a four-point scale (Fully Achieved, Mostly Achieved, Partly Achieved, Not Achieved), and a certified assessor moderates the reports together to issue a final Satisfactory or Unsatisfactory judgement per competency, with Satisfactory required in all areas to pass.

Before trusting the framework, the team checked it two ways: replaying real trainee trajectories through the digital twin to compare trajectory accuracy, and having multiple instructors score the same runs — mixing human and agent runs — to measure how much their ratings agreed.

Why This Matters

Impact on research: The paper argues that academic ATC research has repeatedly simplified the problem to fit available techniques, so results do not transfer to operations. By anchoring evaluation in a regulator-certified curriculum, the framework gives a common, domain-authentic yardstick that can be applied across very different agent approaches, including rules-based systems, evolutionary optimisation, and reinforcement learning. It also produces qualitative expert feedback that can be turned into numerical reward functions, providing a route to close the representational gap.

Real-world applications:

  • Decision support tools that suggest controller actions rather than only flagging conflicts.
  • Trainee controller assessment and standardisation in ATC training colleges, using replayable scenarios.
  • Digital twin assurance: verifying that a simulator faithfully reproduces real exercises before it is used to certify anything.
  • Regulatory assurance of autonomous systems in safety-critical domains, using a licensing-style curriculum as the acceptance standard.

Industry relevance: The work is grounded in NATS, the UK air traffic services provider, with the regulator (the Civil Aviation Authority) underpinning the training requirements and with Project Bluebird as an EPSRC Prosperity Partnership between NATS, The Alan Turing Institute and The University of Exeter. The paper explicitly frames "ensuring separation" as a standard that goes beyond collision prediction, which is directly relevant to any manufacturer or operator seeking to certify automated control assistance.

Future Directions

  • Improve simulation fidelity: Add functional support for a wider range of procedures, services outside controlled airspace, airport departure management, and other ATC concepts, each of which will require further competency mapping.

  • Release an open curriculum: The formative and summative exercises used here cannot be publicly disclosed due to proprietary intellectual property constraints, so the team plans to release an open-source airspace sector of similar fidelity with pre-designed and automatically generated test scenarios.

  • Build expert-validated numerical measures: Develop numerical objectives and reward functions refined against expert feedback, with initial work presented in reference [28], and evaluate agents' emergent behaviour under those measures.

  • Pursue formal assurance discussions: Engage the regulator on establishing more formal frameworks for assurance, and expand testing to further agents, such as the work of Kent et al. [24], when mature.

  • Incorporate probabilistic aircraft modelling (reference [29]) into scenario design to reflect real-world behaviour.

Target Audience

Researchers working on AI for air traffic control or on reinforcement learning and optimisation for safety-critical control; human factors and HCI researchers studying human-machine teaming and expert-in-the-loop evaluation; ATC instructors, training providers, and regulators interested in how AI systems might be assessed against existing licensing standards; and AI assurance specialists looking for a worked example of aligning machine evaluation with a legally regulated professional curriculum.

Authors’ abstract

We present a rigorous, human-in-the-loop evaluation framework for assessing the performance of AI agents on the task of Air Traffic Control, grounded in a regulator-certified simulator-based curriculum used for training and testing real-world trainee controllers. By leveraging legally regulated assessments and involving expert human instructors in the evaluation process, our framework enables a more authentic and domain-accurate measurement of AI performance. This work addresses a critical gap in the existing literature: the frequent misalignment between academic representations of Air Traffic Control and the complexities of the actual operational environment. It also lays the foundations for effective future human-machine teaming paradigms by aligning machine performance with human assessment targets.

Read the original paper