Skip to content
AI.info

Research

Robot-Assisted Group Tours for Blind People

Overview Research area: Human-Computer Interaction / accessibility technology, specifically assistive mobile robots for blind and visually impaired people in social group settings. Technical level: In

Robot-Assisted Group Tours for Blind People
arXiv
2602.04458
Published
2026-02-04
Authors
Yaxin Hu, Masaki Kuribayashi, Allan Wang, Seita Kayukawa, Daisuke Sato, Bilge Mutlu, Hironobu Takagi, Chieko Asakawa

AI summary

Overview

  • Research area: Human-Computer Interaction / accessibility technology, specifically assistive mobile robots for blind and visually impaired people in social group settings.
  • Technical level: Intermediate. The paper combines qualitative HCI study design with a prototype robot system (vision-language model, UWB localization, LiDAR, teleoperation), so familiarity with HCI study methods and basic robotics concepts helps.
  • Scope: A three-phase study (interview study, system design and building, field study) on how an assistive mobile robot can help blind visitors participate in guided museum group tours alongside sighted visitors.

What This Paper Is About

Group interactions normally depend on reading visual cues — who is speaking, where people are standing, when the group moves — which makes group participation hard for blind people. The authors investigate whether a mobile robot can help blind visitors join a guided museum tour that also includes sighted visitors and a human guide. They first interview blind people and museum science communicators to learn the needs and challenges, then build a robot system with five features, and finally test it in a real science museum with mixed-visual tour groups.

Key Contributions

  1. Design requirements: An interview study with blind people (n=5) and museum experts (n=5, science communicators) that documents the structure of tours for blind visitors, the challenges faced by blind visitors and guides, and requirements for a robotic system.
  2. System implementation: A working robot prototype built on the open-source Cabot platform, with five features organized under three design goals — environment awareness, group interaction awareness, and independent navigation — controlled through four buttons on the robot handle.
  3. Empirical findings: A field study in a science museum with blind participants (n=8) and sighted participants (n=8), where each tour group consisted of one guide, one blind participant, and two sighted participants, revealing robot use patterns for group following and tour engagement.
  4. Design implications: Guidance for future robotic systems that support blind people's participation in mixed-visual groups.

Main Findings

  • Tour challenges reported by blind visitors: Three themes emerged from the interviews — difficulty communicating with the tour guide (approaching the guide, hearing the guide in noisy environments, finding the right moment to ask questions), difficulty following the tour (anxiety about moving through a group, falling behind without noticing group movement), and lack of a personalized tour pace (moving on before forming a mental image, less touching than desired).
  • Differences in blind-visitor tours reported by science communicators: Three themes — a focus on the touch experience rather than visual description alone, the need to explain visual information (exhibit shape, color, position, distance; room and nearby-people descriptions), and heavier tour group management, including extra staff to manage flow and "the traffic from sighted visitors," sometimes a three-person team.
  • Robot design insights from interviews: Participants wanted the robot to explain visual, tactile, and environmental information and point out unknown things; to support group turn taking, guide identification, and following; and to support proximity and safety in navigation. Several participants suggested vibration (to avoid interfering with hearing) and light (to signal subtly) instead of loud sounds.
  • Three design goals translated into five robot features: (1) describe the nearby environment using the semantic map and the robot's live camera images (LEFT button); (2) describe the guide's position (RIGHT button); (3) describe nearby people (press and hold RIGHT button); (4) notify the guide (DOWN button); (5) always follow the guide and stop when the guide stops (no user input). The UP button stops all robot speech.
  • System behavior examples: The guide description distinguishes whether the guide faces the user, e.g., "The guide is facing you at 10 o'clock and is three meters away" versus "The guide is three meters away, not facing you." Nearby-people descriptions include counts and distances, such as "Three people are nearby, with one point five, two, and three meters away from the robot."
  • Field study themes: Findings indicated users' sense of safety from the robot's navigational support, concerns in group participation, and preferences for obtaining environmental information. The authors emphasize the importance of a sense of safety, user control and effective robot feedback during navigation, balancing information from the robot versus the guide, and helping users maintain connections with the guide on the tour.
  • Reported participant demographics: Blind interview participants were 49–79 years old (M=57.00, SD=11.19, Women=4, Men=1); four had joined guided museum group tours and one had joined a lecture-style group activity; all five had used similar suitcase-shaped navigation robots. Science communicators were 28–32 years old (M=29.20, SD=1.60, Women=2, Men=3) with 1.5–4 years at the museum (M=2.4, SD=0.97); three had hosted tours or events for blind and low-vision people and two had organized tours for the deaf and hard-of-hearing community.

Methodology in Plain English

The researchers used a three-phase process.

  1. Interview study (needfinding). They recruited five blind people through a mailing list maintained by their institution and five museum science communicators through the museum's management team. Blind participants were interviewed online for 30 minutes each about past tour experiences and their experience with navigation robots. Science communicators were interviewed in a museum conference room for one hour each, and after the questions they watched a two-minute video of a suitcase-shaped assistive robot navigating in a shopping mall, then generated ideas for using the robot in tours. All interviews were audio recorded, and the team analyzed them with reflective thematic analysis, using two coders and affinity diagramming.
  2. System design and building. Insights from the interviews were condensed into three design goals and five robot features. The system used the open-source Cabot robot platform, a phone on the user side running a User App, a phone worn by the guide running a Guide App, a bone-conduction headset for audio, and a robot handle with four mechanical buttons and three tactile protrusions housing a vibrator. Surrounding environment descriptions were generated by GPT-4o using three RGB images (left, front, right) plus exhibit information based on the robot's location and orientation. Guide position came from UWB signals broadcast by the guide's phone. Nearby people were detected with the robot's cameras and a local detection model. The robot navigated by teleoperation while using LiDAR for collision avoidance. The robot provided a WiFi network for WebSocket communication with the phone apps and used LTE to reach the LLM server.
  3. Field study. In a science museum, they ran tours with eight blind participants and eight sighted participants, with each group containing one guide, one blind participant, and two sighted participants. Each blind participant had a training session with the guide before the tour. After the tours, they interviewed blind participants, sighted participants, and tour guides.

All study protocols and materials were approved by Waseda University's institutional review board (IRB).

Why This Matters

This work pushes assistive robotics past point-to-point navigation and individual exploration into social, multi-person settings where a blind person must coordinate with a human guide and other visitors. It shows concretely how a robot can share attention and information across three stakeholders (blind visitor, guide, sighted visitors) rather than serving one user alone.

Real-world applications:

  • Museums and cultural institutions that offer regular tours which are typically not accessible to blind visitors, alongside less frequent special tours.
  • Public guided experiences such as gardens, art galleries, theaters, and other guided leisure contexts where group following and guide contact matter.
  • Dynamic public spaces like shopping malls or unfamiliar buildings, where preserving appropriate social distance and alerting the user to nearby people is useful.
  • Human-robot communication support, including signaling turn-taking in group conversation and giving a blind user a private channel to request information or notify a guide.

Industry relevance:

  • Assistive robotics and mobility hardware: the work builds on an open-source platform (Cabot) and records user preferences for wheeled robots (quieter, more stable, smoother) versus quadrupeds, which can handle stairs and uneven terrain.
  • Vision-language models in accessibility: the prototype uses GPT-4o to narrate the surroundings from three RGB images plus map-based exhibit metadata, an example of pairing foundation models with spatial context.
  • Indoor localization and multisensory feedback: UWB for guide tracking and vibration or light alerts instead of loud audio point to design constraints that matter for any wearable or robot product used in public, quiet environments.

Future Directions

  • Rebalance robot and guide information. The findings call out the need to balance what the robot says with what the human guide says, so the two do not compete for the user's attention or duplicate content.
  • Strengthen user control and feedback during navigation. Users wanted to feel in control and to receive effective feedback while being guided, which raises questions about how much autonomy the robot should have versus teleoperation by an experimenter as in this prototype.
  • Design for the whole group, not just the blind user. The paper asks how robot-assisted participation affects other group members, so future work could examine effects on guides and sighted visitors and support more than one blind participant per tour.
  • Extend beyond the museum tour prototype. The authors present design implications for future robotic systems supporting mixed-visual group participation broadly; the specific settings, durations, and configurations tested here leave open how the approach generalizes.

Target Audience

Researchers and practitioners in human-robot interaction, accessibility technology, and assistive robotics; museum and cultural-institution staff who design or run accessible tours; and designers building vision-language-model or localization-based assistance features for blind and low-vision users. It is most useful to readers who already understand basic HCI study methods and want a concrete example of designing and evaluating a social assistive robot in a real public space. Readers looking for quantitative performance metrics or benchmark results will not find them reported in the provided content, which centers on qualitative findings.

Authors’ abstract

Group interactions are essential to social functioning, yet effective engagement relies on the ability to recognize and interpret visual cues, making such engagement a significant challenge for blind people. In this paper, we investigate how a mobile robot can support group interactions for blind people. We used the scenario of a guided tour with mixed-visual groups involving blind and sighted visitors. Based on insights from an interview study with blind people (n=5) and museum experts (n=5), we designed and prototyped a robotic system that supported blind visitors to join group tours. We conducted a field study in a science museum where each blind participant (n=8) joined a group tour with one guide and two sighted participants (n=8). Findings indicated users' sense of safety from the robot's navigational support, concerns in the group participation, and preferences for obtaining environmental information. We present design implications for future robotic systems to support blind people's mixed-visual group participation.

Read the original paper