Skip to content
AI.info

Research

Reimagining the Augmented Reality Accessibility Ecosystem for Deaf Students: Service Provider Perspectives in Experiential Learning

Overview Research area: Human-Computer Interaction, specifically assistive and accessible computing — augmented reality (AR) communication access for Deaf and Hard of Hearing (DHH) students in experie

arXiv
2607.21289
Published
2026-07-23
Authors
Roshan Mathew, Roshan Peiris

AI summary

Overview

Research area: Human-Computer Interaction, specifically assistive and accessible computing — augmented reality (AR) communication access for Deaf and Hard of Hearing (DHH) students in experiential (hands-on) learning, examined from the perspective of the service providers rather than the student end user.

Technical level: Intermediate. The paper describes an AR smart-glasses platform and a multi-camera streaming pipeline, but its core content is a qualitative, expert-based evaluation using reflexive thematic analysis. No statistical modeling or quantitative benchmarks are involved.

Scope: A formative, expert-based evaluation of ARRAE (Augmented Reality Real-Time Access for Education) in a simulated laboratory, comparing in-person, traditional remote, and AR-mediated communication access through the experiences of one instructor, one ASL interpreter, and one real-time captioner.

What This Paper Is About

In hands-on learning environments such as chemistry laboratories, DHH students face a "split-attention effect": they must divide their focus between the experimental task (where looking away can be a safety hazard) and the interpreter, captioner, or instructor. AR smart glasses have been proposed as a way to overlay communication access directly into the student's field of view, but prior evaluations have almost exclusively examined the student's user experience — leaving unknown how such systems affect the instructors, interpreters, and captioners who deliver access. This paper evaluates ARRAE from those providers' perspectives, asking how AR reconfigures accessibility labor, communication, attention, and awareness across all participants.

Key Contributions

  1. A qualitative account of how AR-mediated communication access reshapes accessibility labor and interactional practices among instructors, interpreters, and captioners during hands-on laboratory tasks, compared against in-person and conventional remote access.
  2. An articulation of key usability, workflow, and design considerations that emerge when deploying AR-mediated communication access in experiential higher education settings.
  3. A reframing of communication access as distributed awareness — a shared, multi-user awareness environment rather than an individual assistive tool — supported by a comparison table of modalities across six themes.
  4. An argument that AR-mediated access is best understood as part of a broader hybrid ecosystem rather than a standalone replacement for existing modalities.

Main Findings

  • Familiarity and reliability: In-person access was consistently described as the most natural and reliable. The instructor said in-person interpreting "feels the most familiar"; interpreters called it "least cognitively demanding"; captioners called in-person captioning "effortless and highly familiar." But in-person access limited visibility and introduced spatial constraints — the instructor could not see the student's face or workstation, and noted that "having another body in the room is always difficult in a science lab."

  • Visual awareness and task coordination: Remote and ARRAE conditions improved instructor awareness. The instructor reported being able to see the student's face, a limited view of the workstation, and the interpreter, which helped with pacing instructions. ARRAE's multi-camera views further supported this ("Multiple camera views is very helpful..."). Captioners also valued multiple camera angles for context. However, visibility remained fragile: occlusion from materials disrupted communication in both remote and ARRAE setups.

  • Attention integration and interactional flow: Remote setups required attention shifts or physical movement — students had to "move between the supply desk and the fume hood," and Zoom captions required looking away. ARRAE integrated communication into the student's field of view, letting students work while watching the interpreter or reading captions. Interpreters reported reduced reliance on gaze alignment. However, ARRAE introduced coordination challenges: instructors found it difficult to initiate interaction, requiring coordination with the interpreter and camera positioning.

  • Instructor awareness and control of communication: In remote captioning, instructors valued seeing the captions to gauge their own pacing. In ARRAE, this awareness was reduced — the instructor wished they could see the captions to know whether instructions were conveyed correctly and whether to slow down. In-person setups provided indirect awareness through environmental cues such as seeing or hearing the captionist type.

  • Usability and workflow constraints: Remote setups suffered from audio capture issues, setup complexity, and platform limitations (muting everyone except the instructor, granting manual captioning ability). ARRAE improved audio quality, contextual awareness, and interface ergonomics — the captioner described setup and delivery as "pretty seamless" with much clearer audio than Zoom. But ARRAE introduced system instability ("freezing of the video feed"), limited interface flexibility, and hardware constraints.

  • Persistence of information: In-person captioning supported task continuity through stable text displays that students could refer back to. Zoom captions were described as transient, disappearing quickly. ARRAE improved persistence by keeping captions visible in the student's field of view, letting students refer back while gathering materials. This was limited by the space available in the captioning container seen within the AR view.

  • Overall trade-offs: In-person access offers reliability but limits visibility and scalability; remote approaches improve awareness but introduce fragmentation and technical challenges; AR-mediated access integrates communication into ongoing activity and enhances contextual awareness but currently limits instructor oversight and presents usability constraints.

Methodology in Plain English

The researchers built ARRAE, a platform that connects support personnel to DHH students wearing optical see-through smart glasses (Vuzix Blade). Students received either a live video feed of an interpreter or a scrolling text stream of captions directly in their line of sight, following the spatial contiguity principle. Support personnel used a web-based dashboard fed by three Pan-Tilt-Zoom cameras in the lab — a Workstation Camera (workbench view), an Observation Camera (wide-angle view of the student's signing space), and an Instructor Camera (instructor view) — streamed via WebRTC with sub-second latency. Interpreters broadcast their video to the student's glasses; captioners worked through a view integrated with a TypeWell ingestion server; the instructor's dashboard aggregated all active feeds.

The evaluation used a simulated laboratory at Location A (a workstation with a biosafety cabinet, instrumented with an OBSBOT Tail Air 4K PTZ camera for the workbench, an overhead Tenveo 10x Zoom PTZ camera for the signing space, and an OBSBOT Tiny 4K webcam for the instructor) and a Remote Support Hub at Location B (a laptop and external monitor). The broader study involved 12 deaf student participants — six in the interpreting study and six in the captioning study — whose perspectives are reported separately. This paper focused on three expert participants: one STEM instructor (10–15 years of experience), one certified ASL interpreter (3–5 years of VRI experience), and one professional captioner (5–10 years of real-time captioning experience). The instructor participated in both studies, which were run separately; the interpreter only in the interpreting study and the captioner only in the captioning study.

Each study included six sessions with six different deaf student participants, segregated by preferred primary access modality. In each session, the student completed three hands-on laboratory tasks (making artificial snow using an instant snow polymer) guided by spoken English instructions, experiencing all three access conditions: in-person access, remote access (traditional VRI/captioning via tablet), and AR-mediated access (ARRAE). Condition order and task order were counterbalanced; sessions were moderated by deaf researchers, and students received smart-glasses training beforehand. After each session, the instructor and relevant access provider completed open-ended surveys; after all six sessions in each study, a brief follow-up semi-structured interview was conducted and the moderator recorded interview notes. These data were analyzed inductively using reflexive thematic analysis, focusing on recurring patterns rather than thematic saturation or generalizable findings.

Why This Matters

For research, the paper shifts the unit of analysis for AR accessibility from the individual end user to the multi-stakeholder ecosystem. It argues that a system which reduces student cognitive load but increases interpreter burden or disconnects the instructor from their instructional rhythm causes the whole ecosystem to fail — and it introduces "distributed awareness" as a framing for evaluating such systems.

Real-world applications:

  • Laboratory and clinical training: Hands-on science, medical, and technical instruction where students cannot look away from a task, a specimen, or a safety-critical procedure.
  • Remote interpreting and captioning services: Informing how VRI and real-time captioning providers are equipped with camera views, dashboards, and workflows that preserve context.
  • Classroom accessibility provisioning in higher education: Guiding institutions deciding how to combine in-person, remote, and AR support rather than replacing one with another.
  • Accessible AR/AR hardware design: Feeding requirements such as subtitle persistence, container sizing within the AR field of view, and attention-signaling mechanisms back into device and interface design.

Industry relevance: the findings matter to vendors of smart glasses and AR headsets, remote captioning and interpreting platforms, and PTZ camera/streaming suppliers, as well as to post-secondary disability services offices. The paper argues AR-mediated access is most effective as part of a broader hybrid ecosystem, which is directly relevant to procurement and deployment decisions, not just to prototype development.

Future Directions

  1. Replication with more professionals: Multiple instructors, interpreters, and captioners in longitudinal studies are needed to test whether observed patterns hold over time and how increasing familiarity reshapes coordination demands, usability, and instructional fit.
  2. Repairing fractured interactional feedback: Addressing reduced access to non-verbal and backchanneling cues used to gauge comprehension, pacing, and communicative alignment — including better attention-signaling mechanisms for initiating interaction.
  3. Role-specific design refinement: Role-specific visual views, shared awareness features so instructors can see captions, improved persistence and review of captioning content, more flexible interface controls, and more dependable audio and video performance.
  4. Affective and multi-element speech presentation: Exploring how additional elements of speech, such as affective captions, can be effectively delivered from a service-provider perspective and what impact they have on user experience, across diverse experiential tasks.

Target Audience

This paper is most valuable to HCI and accessibility researchers working on AR, assistive technology, and multi-user interaction; designers and engineers building AR smart-glasses communication tools; interpreting and captioning professionals and the agencies that deploy VRI and real-time captioning; disability services and instructional staff in post-secondary STEM education; and AR hardware and remote-access platform vendors seeking requirements grounded in service-provider workflows.

Authors’ abstract

In experiential learning environments, Deaf and hard of hearing (DHH) students often experience ``split attention,'' dividing their focus among tasks, instructors, and access providers. Augmented reality (AR) has been proposed as a means to centralize communication access within the student's field of view; however, little is known about how such systems affect the instructors, interpreters, and captioners who support access in these settings. We present a formative, expert-based evaluation of ARRAE, an AR-mediated communication access ecosystem, examining the experiences of an instructor, an American Sign Language (ASL) interpreter, and a real-time captioner in a simulated laboratory environment. Comparing in-person, traditional remote, and AR-mediated access, our findings suggest that AR reconfigures accessibility labor and interactional practices by redistributing communication, attention, and awareness across participants. While AR-mediated access supported integrated, in-situ communication and enhanced contextual awareness, it also altered interactional feedback mechanisms, such as gaze coordination and pacing, that support shared understanding during instruction. These observations highlight trade-offs between access, control, and coordination, and suggest that AR-mediated systems may be most effective when considered as part of a broader, hybrid ecosystem rather than as standalone replacements for existing modalities. This work contributes initial insights into how AR reshapes communication access in experiential learning and informs the design of systems that support coordinated, multi-user interaction in hands-on educational contexts.

Read the original paper