Skip to content
AI.info

Research

Integrating Virtual Reality and Large Language Models for Team-Based Non-Technical Skills Training and Evaluation in the Operating Room

Overview Research area: Human-Computer Interaction applied to surgical education — specifically multi-user virtual reality (VR) simulation combined with large language model (LLM) analytics for team-b

Integrating Virtual Reality and Large Language Models for Team-Based Non-Technical Skills Training and Evaluation in the Operating Room
arXiv
2601.13406
Published
2026-01-19
Authors
Jacob Barker, Doga Demirel, Cullen Jackson, Anna Johansson, Robbin Miraglia, Darian Hoagland, Stephanie B. Jones, John Mitchell, Daniel B. Jones, Suvranu De

AI summary

Overview

Research area: Human-Computer Interaction applied to surgical education — specifically multi-user virtual reality (VR) simulation combined with large language model (LLM) analytics for team-based non-technical skills (NTS) training.

Technical level: Intermediate. The concepts are accessible to a general reader, but the work assumes familiarity with surgical education frameworks (NOTSS, ACS/APDS curriculum), VR simulation, and LLM-based text analysis.

Scope: The paper introduces VORTeX, a multi-user VR operating room platform that pairs immersive team scenarios with LLM-driven analysis of team dialogue to teach and objectively evaluate communication, decision-making, teamwork, and leadership.

What This Paper Is About

Surgical safety depends heavily on teamwork and communication, yet structured training for these non-technical skills lags far behind technical skills simulation. The ACS/APDS Phase III Team-Based Skills Curriculum explicitly calls for scalable tools that can both teach and objectively assess these competencies during laparoscopic emergencies.

This paper's goal is to meet that call by building a system that lets surgical teams practice together in immersive VR while an LLM automatically analyzes their spoken dialogue to classify team behaviors and map communication structure, rather than relying solely on subjective human rating.

Key Contributions

  1. The VORTeX platform — a multi-user VR environment that combines immersive operating room team simulation with integrated LLM analytics for training and evaluating communication, decision-making, teamwork, and leadership.
  2. A NOTSS-grounded analytics pipeline — team dialogue is analyzed using structured prompts derived from the Non-Technical Skills for Surgeons (NOTSS) framework, enabling automated classification of observable behaviors.
  3. Directed interaction graphs — a representation generated from the dialogue analysis that quantifies communication structure and hierarchy within the team, not just the presence of behaviors.
  4. Two implemented laparoscopic emergency scenarios — pneumothorax and intra-abdominal bleeding — designed to elicit realistic stress and collaboration, with pilot sessions conducted with surgical professionals at the 2024 SAGES conference.

Main Findings

  • Positive pilot reception: The twelve surgical professionals who completed pilot sessions at the 2024 SAGES conference rated VORTeX as intuitive, immersive, and valuable for developing teamwork and communication. The abstract does not report the specific rating instrument or scoring scale used.
  • Interpretable communication networks: The LLM consistently produced communication networks that were readable and that reflected the operative hierarchies the authors expected to see. The abstract gives no quantitative accuracy or agreement figures.
  • Distinct role patterns emerged: Surgeons appeared as central integrators within the networks, nurses as initiators, and anesthesiologists as balanced intermediaries — suggesting the analytics capture meaningful role-based structure in team dialogue.
  • Scenario design elicited the intended behavior: The two laparoscopic emergency scenarios (pneumothorax and intra-abdominal bleeding) were reported to prompt realistic stress and collaboration among participants.
  • A claimed scalable, privacy-compliant framework: The authors position VORTeX as supporting objective assessment and automated, data-informed debriefing across distributed training environments. These are stated as design properties; the abstract reports no comparative benchmarks against existing assessment methods.

Methodology in Plain English

The researchers built a shared virtual operating room in which multiple participants can join the same simulated scenario at once. Two laparoscopic emergency cases were authored to create pressure and force the team to communicate and coordinate.

While the team works through the scenario, their spoken dialogue is captured. That dialogue is then fed to a large language model along with structured prompts written around the NOTSS framework — a recognized taxonomy of surgical non-technical skills. The prompts direct the model to classify what kinds of behaviors are occurring in the conversation. From those classifications, the system builds directed interaction graphs, which show who is talking to whom and how often, making the shape and hierarchy of team communication visible as a network rather than as a raw transcript.

To get an initial read on whether the approach works in practice, the team ran pilot sessions with surgical professionals at a major surgical conference and collected their impressions of the experience.

Why This Matters

Research impact: The work demonstrates a way to make non-technical skill assessment more objective and reproducible. Instead of depending entirely on expert raters watching recordings, the LLM generates a structured, inspectable representation of team communication. It also shows that VR simulation and language-model analytics can be combined into a single training-and-assessment loop, which is a different framing than treating simulation and evaluation as separate activities.

Real-world applications:

  • Surgical residency and team training programs that need scalable ways to teach and evaluate communication during laparoscopic emergencies.
  • Debriefing after simulation sessions, where automatically generated communication networks could give instructors concrete, data-informed talking points rather than relying on memory.
  • Distributed or remote training, where teams in different locations could be assessed under a common, automated standard.
  • Competency documentation, potentially feeding structured evidence into the assessment requirements of curricula such as the ACS/APDS Phase III Team-Based Skills Curriculum.

Industry relevance: The combination of consumer-accessible VR hardware with LLM analytics points toward a commercially plausible model for simulation-based clinical education — one where software, rather than scarce expert faculty time, handles a substantial share of the assessment load. The emphasis on privacy compliance also signals awareness of the institutional and regulatory constraints that govern handling of clinical training data.

Future Directions

  • Validation beyond a pilot. The abstract reports only impression ratings from twelve professionals at one conference. Establishing whether the LLM's classifications agree with expert human raters, and whether they predict real operative performance, remains open.
  • Scaling the scenario library. Only two laparoscopic emergency scenarios are described. Coverage of a broader range of emergencies, specialties, and team compositions would be needed for general use.
  • Cross-site and longitudinal study. Whether automated debriefing actually improves team performance over repeated sessions across distributed training environments is a natural next question the framework sets up but does not answer here.
  • Generalization and bias in the analytics. The abstract does not address how the NOTSS-derived prompts behave across different accents, communication styles, or institutions, nor how the LLM handles ambiguous or overlapping speech — areas that would need scrutiny before deployment.

Target Audience

This paper is most useful to surgical educators and simulation center directors looking for scalable non-technical skills assessment; human-computer interaction researchers working on VR collaboration or LLM-based behavioral analytics; medical education researchers interested in automated competency assessment; and clinical curriculum designers responding to framework requirements such as the ACS/APDS Phase III Team-Based Skills Curriculum. Readers seeking rigorous quantitative validation will find only pilot-level evidence here, as the abstract reports no performance metrics beyond participant impressions.

Authors’ abstract

Although effective teamwork and communication are critical to surgical safety, structured training for non-technical skills (NTS) remains limited compared with technical simulation. The ACS/APDS Phase III Team-Based Skills Curriculum calls for scalable tools that both teach and objectively assess these competencies during laparoscopic emergencies. We introduce the Virtual Operating Room Team Experience (VORTeX), a multi-user virtual reality (VR) platform that integrates immersive team simulation with large language model (LLM) analytics to train and evaluate communication, decision-making, teamwork, and leadership. Team dialogue is analyzed using structured prompts derived from the Non-Technical Skills for Surgeons (NOTSS) framework, enabling automated classification of behaviors and generation of directed interaction graphs that quantify communication structure and hierarchy. Two laparoscopic emergency scenarios, pneumothorax and intra-abdominal bleeding, were implemented to elicit realistic stress and collaboration. Twelve surgical professionals completed pilot sessions at the 2024 SAGES conference, rating VORTeX as intuitive, immersive, and valuable for developing teamwork and communication. The LLM consistently produced interpretable communication networks reflecting expected operative hierarchies, with surgeons as central integrators, nurses as initiators, and anesthesiologists as balanced intermediaries. By integrating immersive VR with LLM-driven behavioral analytics, VORTeX provides a scalable, privacy-compliant framework for objective assessment and automated, data-informed debriefing across distributed training environments.

Read the original paper