Research
Beyond Static Instruction: A Multi-agent AI Framework for Adaptive Augmented Reality Robot Training
Overview Research area: Human-robot interaction (HRI), augmented reality (AR) interfaces for industrial robot training, applied large language model (LLM) multi-agent systems, and instructional design

- arXiv
- 2603.00016
- Published
- 2026-01-31
- Authors
- Nicolas Leins, Jana Gonnermann-Müller, Malte Teichmann, Sebastian Pokutta
AI summary
Overview
- Research area: Human-robot interaction (HRI), augmented reality (AR) interfaces for industrial robot training, applied large language model (LLM) multi-agent systems, and instructional design.
- Technical level: Intermediate. Readers need some familiarity with AR/VR hardware, ROS-based robot control, and the general capabilities and limitations of LLM agents, but the paper explains its architecture in plain terms.
- Scope: The paper presents a fully implemented AR training application for a Universal Robots UR5e arm, reports a 36-participant usability evaluation of that baseline, and proposes (but does not yet implement or evaluate) a multi-agent LLM framework intended to make the training environment adapt dynamically to individual learners.
What This Paper Is About
Industrial robot training still relies on static resources — manuals, videos, and rigid step-by-step instruction — and even AR interfaces that visualize robot motion in 3D tend to show the same content to every user regardless of skill, stress, or spatial ability. The authors built an AR robot-training application and tested it with novices, finding that although usability was high overall, learners with lower spatial ability, lower affinity for technology, or less robotics experience reported more cognitive load and lower usability. To close that gap, the paper proposes a conceptual multi-agent architecture in which specialized LLM agents process multimodal sensor data and pedagogical rules to adapt the AR environment in real time.
Key Contributions
-
A fully implemented, open-source AR robot training application. Built in Unity (2022.3) and deployed on the Meta Quest 3 video see-through head-mounted display, it uses bare-hand tracking to replicate the functionality of an industry teach pendant for a Universal Robots UR5e arm, with support for joint-based and translational Tool Center Point (TCP) movement plus waypoint programming. The code is publicly available in a repository.
-
A preliminary user evaluation of the baseline interface with N = 36 participants. The study assessed System Usability Scale (SUS), extraneous cognitive load (ECL), Mental Rotation Test (MRT), Affinity for Technology Interaction (ATI), and prior experience in robotics (ER) after a standardized three-part learning task.
-
A conceptual multi-agent AI framework for adaptive training. The architecture has three layers — a deterministic Input Layer for sensor preprocessing, an LLM-driven Reasoning Layer for pedagogical decision-making, and an Output Layer of specialized agents that execute concrete interventions in the AR application.
-
An explicit treatment of ethical considerations and limitations, including three strategies for handling biometric data privacy and an acknowledgment of LLM non-determinism as an unresolved risk.
Main Findings
-
High baseline usability: The AR system achieved an overall mean SUS score of 82.6 (SD = 14.1), described in the paper as corresponding to a very high rating, with a low overall ECL of M = 1.70 (SD = 0.7). The authors read this as evidence that the learning material and interface design were appropriate for supporting the task.
-
Wide variation in completion time: Task duration ranged from 14 to 33 minutes (M = 23.1, SD = 4.8), revealing substantial differences in how long participants needed.
-
Prior experience mattered: High-experience users reported lower ECL (M = 1.56) and higher SUS (M = 89.2) than low-experience users (ECL: M = 1.78, SUS: M = 78.8).
-
Spatial ability mattered: High-MRT users reported lower ECL (M = 1.55) and higher SUS (M = 85.1), while low-MRT users showed higher ECL (M = 1.84) and lower SUS (M = 80.3).
-
Technology affinity produced the largest usability gap: Users with low ATI rated SUS nearly 13 points lower (M = 75.8) than high-ATI users (M = 88), and reported higher ECL (M = 1.85 vs. M = 1.58).
-
The pattern supports the "one size does not fit all" argument: Participants with lower spatial ability, lower ATI, or less experience perceived the static interface as more burdensome and less usable, which the authors use as motivation for the proposed adaptive framework.
-
The adaptive framework itself is not yet evaluated: The paper is explicit that the multi-agent integration is a proposed architecture for future development; only the baseline AR application was implemented and tested. No learning-outcome comparison between static and adaptive conditions is reported.
Methodology in Plain English
The work proceeds in two stages. First, the team built the AR application: a Unity program running on a Meta Quest 3 headset that communicates bidirectionally with a UR5e robot through ROS2 using a Unity-ROS-TCP connector. The app mirrors the robot's joint states in real time, publishes URScript commands to control the physical arm, and anchors virtual content to the robot's base. Users interact by pressing virtual buttons with their index fingers — no handheld controllers. To help learners, the interface shows the robot's coordinate system at the TCP, spatial waypoint markers connected by path lines, and dynamic tooltips tied to the current learning step.
Second, they designed a standardized three-topic learning task for novices: (1) basic joint control to reach a target pose, (2) linear TCP displacement in Cartesian space to reach the same pose, and (3) programming a full pick-and-place sequence with the mounted OnRobot RG2 gripper, including waypoints and intermediate support positions to avoid collisions. Thirty-six participants (27 men, nine women, average age 27, SD = 6.7) completed the task and then filled out questionnaires covering system usability (SUS), extraneous cognitive load (the naive rating scale from Klepsch et al.), mental rotation ability (MRT), affinity for technology interaction (ATI), and prior robotics experience (ER). Participants were split post hoc into "High" and "Low" groups for MRT, ATI, and ER based on the sample mean.
Third, based on these results, the authors designed — on paper — a layered agent architecture. The Input Layer contains deterministic modules (Voice Analyzer, Progress Analyzer, Robot Data Analyzer, Physiological Analyzer) that convert raw sensor streams into structured semantic events, such as turning eye-tracking coordinates into the statement "user is fixating on the gripper," specifically to prevent early-stage hallucination. The Reasoning Layer splits understanding from deciding: an Assessment Agent summarizes the user's situation — for example, flagging frustration at step four — while a Teacher Agent uses that summary plus a knowledge base of instructional content and pedagogical rules to choose an intervention. The Output Layer contains a Tutor Agent (empathetic spoken text via a virtual avatar), a Visualization Agent (for example, generating an extra arrow to guide a user struggling with mental rotation, with position, scale, and color parameters), and an Instruction Agent (rewriting complex explanations into simpler language). Reasoning Layer agents use high temperature settings for creative reasoning; Output Layer agents use low temperature settings and must adhere strictly to predefined JSON schemas that the AR application parses to trigger local functions. The design is grounded in Mayer's Multimedia Principles and cognitive load theory.
Why This Matters
Impact on research. The paper connects two literatures that are often separate: AR interfaces for industrial robot training and multi-agent LLM orchestration. Its distinctive move is to split the "teacher" into specialized agents across a deterministic-to-generative spectrum, arguing that deterministic preprocessing and schema-constrained outputs are what make LLM-driven adaptation safe enough for a real training context. It also contributes a documented baseline dataset of learner-characteristic differences (MRT, ATI, ER) that future adaptive systems can be measured against.
Real-world applications:
- Industrial onboarding and vocational training, where new operators must learn to control and program robot arms that are increasingly common on factory floors.
- Adaptive AR/VR training platforms in domains beyond robotics where spatial reasoning and procedural learning matter, such as maintenance, assembly, or medical device operation.
- Operator assistance in small and mid-sized manufacturers, where a single AR system may need to serve users with widely different prior experience.
- Human-robot collaboration settings where a system must infer operator stress or confusion from voice, physiology, and robot data and respond without interrupting workflow.
Industry relevance. The bottleneck in industrial robotics is described as shifting from hardware capability to the human operator's ability to interact with it efficiently. Static instructions and 2D interfaces add cognitive load. A system that fades scaffolding as a novice gains confidence and reinstates it when a user struggles could shorten ramp-up time and reduce training burden — a direct operational concern for manufacturers deploying robotic arms.
Future Directions
- Integration and comparative evaluation. The authors state that future work will integrate the multi-agent framework into the AR application and evaluate its impact on learning outcomes in a comparative study. No such study is reported here.
- Reliability of LLM-driven pedagogy. The paper flags the non-deterministic nature of LLMs as a remaining challenge that may produce inconsistent pedagogical decisions, and acknowledges that architectural checkpoints (such as the Teacher Agent validating the Assessment Agent's output) cannot fully eliminate unpredictability. Future work must rigorously evaluate system reliability.
- Extending the sensor and module set. The Input Layer is described as designed to be easily extended with additional modules for other data sources if a use case requires it, leaving open which modalities actually improve adaptation.
- Privacy-preserving deployment. The architecture is optimized for small, locally hosted models so that sensitive biometric data does not leave the user's environment. Whether locally hosted models can match the reasoning quality of larger remote models is an open question.
Target Audience
This paper is most useful to researchers and practitioners working at the intersection of AR/VR training systems, human-robot interaction, and LLM-based agent architectures — particularly those interested in how to structure generative AI so that it is grounded in sensor data rather than free-form. It also suits instructional designers and learning scientists who want a concrete example of how cognitive load theory and multimedia principles are being operationalized for adaptive training. Industrial engineers and training managers evaluating AR for robot onboarding will find the baseline usability numbers and the learner-difference analysis directly relevant, while readers looking for a validated adaptive system should note that the multi-agent portion remains a proposed design.
Authors’ abstract
Augmented Reality (AR) offers powerful visualization capabilities for industrial robot training, yet current interfaces remain predominantly static, failing to account for learners' diverse cognitive profiles. In this paper, we present an AR application for robot training and propose a multi-agent AI framework for future integration that bridges the gap between static visualization and pedagogical intelligence. We report on the evaluation of the baseline AR interface with 36 participants performing a robotic pick-and-place task. While overall usability was high, notable disparities in task duration and learner characteristics highlighted the necessity for dynamic adaptation. To address this, we propose a multi-agent framework that orchestrates multiple components to perform complex preprocessing of multimodal inputs (e.g., voice, physiology, robot data) and adapt the AR application to the learner's needs. By utilizing autonomous Large Language Model (LLM) agents, the proposed system would dynamically adapt the learning environment based on advanced LLM reasoning in real-time.