Research
Explainable OOHRI: Communicating Robot Capabilities and Limitations as Augmented Reality Affordances
Overview Research area: Human-Computer Interaction, specifically explainable human-robot interaction (HRI) and augmented reality interfaces for robotics. Technical level: Advanced. The work combines A

- arXiv
- 2601.14587
- Published
- 2026-01-21
- Authors
- Lauren W. Wang, Mohamed Kari, Parastoo Abtahi
AI summary
Overview
Research area: Human-Computer Interaction, specifically explainable human-robot interaction (HRI) and augmented reality interfaces for robotics.
Technical level: Advanced. The work combines AR interface design, vision-language models, robot manipulation, and a human-subjects evaluation.
Scope: The paper proposes and evaluates X-OOHRI, an AR system that visualizes what a robot can and cannot do, allowing people to give the robot object-oriented commands and to intervene when failure is likely.
What This Paper Is About
Robots are often opaque to the people working alongside them: users get little sense of what the robot is currently capable of, where its limits lie, or why it might fail. This makes it hard to give personalized instructions or to step in and help when a task is about to go wrong. The authors build an AR interface that surfaces robot action possibilities and constraints directly on the objects and virtual counterparts in the scene, so that capabilities and limitations become visible rather than hidden inside the system.
Key Contributions
-
The X-OOHRI interface concept. An augmented reality design that communicates robot action possibilities and constraints using visual signifiers, radial menus, color coding, and explanation tags.
-
An object-oriented representation pipeline. Object properties and robot limits are encoded into object-oriented structures using a vision-language model, which supports generating explanations on the fly.
-
Direct manipulation of virtual twins. Users can act on virtual duplicates of objects that are spatially aligned within a simulated environment, connecting the AR layer to the physical scene.
-
An end-to-end system integrated with a physical robot. The authors demonstrate the full pipeline across use cases spanning low-level pick-and-place operations through to higher-level instructions, and evaluate it with a user study.
Main Findings
-
Object-oriented commanding works: Participants in the user study were able to effectively issue object-oriented commands through the system.
-
Better mental models of limitations: Participants developed accurate mental models of what the robot could and could not do.
-
Mixed-initiative interaction emerges: Participants engaged in mixed-initiative resolution, meaning they actively participated in resolving situations rather than passively watching the robot.
-
Use-case breadth demonstrated: The integrated system was shown to support scenarios from low-level pick-and-place up to high-level instructions.
The abstract reports these outcomes qualitatively; it does not include study size, quantitative measures, or statistical comparisons, so no numerical results can be reported here.
Methodology in Plain English
The researchers designed an AR layer that sits over a robot's working environment. Rather than leaving the robot's reasoning invisible, the interface draws signifiers, radial menus, color coding, and text explanation tags onto objects so that a person can see what actions are available and where the robot's constraints are.
Underneath the interface, object properties and robot limits are turned into object-oriented structures with the help of a vision-language model. That structure is what lets the system produce explanations on the fly, instead of relying on pre-written text. The AR layer also shows virtual twins of objects — virtual copies positioned to line up with the real scene inside a simulated environment — which the user can manipulate directly.
The authors connected this whole pipeline to an actual physical robot, then ran a user study to see how people used it.
Why This Matters
Research impact. The work connects explainable AI with HRI and AR, arguing that explanations are most useful when they are spatially attached to the objects and actions they describe, and when they support the user in acting — not just understanding.
Real-world applications
- Home and assistive robotics: a person could see which objects a helper robot can grasp and which are beyond it, and adjust what they ask for.
- Warehouse and manufacturing: operators supervising pick-and-place robots could spot upcoming failures and intervene before a task goes wrong.
- Remote teleoperation and supervision: an operator at a distance could get capability and constraint information overlaid on the scene rather than inferring it.
- Training and onboarding: new users could build an accurate sense of a robot's limits faster, which the abstract claims the study participants achieved.
Industry relevance. As robots move out of cages and into shared spaces with people, interfaces that make robot limitations legible become a practical requirement for safety, trust, and efficient collaboration. Systems built on vision-language models also point toward interfaces that can be adapted to new objects and tasks without hand-authoring explanations for each one.
Future Directions
- Generalizing the representation: how well the vision-language-model-driven encoding of object properties and robot limits holds up across new objects, tasks, and robot platforms not covered by the demonstrated use cases.
- Longitudinal and quantitative evaluation: the abstract reports a study showing that participants commanded effectively, built accurate mental models, and engaged in mixed-initiative resolution, but longer-term effects on trust, reliance, and skill are open questions.
- From simulation to messy reality: understanding how the spatially aligned virtual twins and explanation tags behave under real-world perception errors, clutter, and occlusion.
- Adapting explanations to the user: whether the signifiers, color coding, and tags can be tuned to different levels of user expertise, and how users should be involved in resolving failures at scale.
Target Audience
HRI and explainable-AI researchers; AR/XR developers building industrial or assistive applications; robotics engineers working on manipulation and human-robot collaboration; and human-factors or UX practitioners interested in how capability and limitation information should be presented to people working alongside robots.
Authors’ abstract
Human interaction is essential for issuing personalized instructions and assisting robots when failure is likely. However, robots remain largely black boxes, offering users little insight into their evolving capabilities and limitations. To address this gap, we present explainable object-oriented HRI (X-OOHRI), an augmented reality (AR) interface that conveys robot action possibilities and constraints through visual signifiers, radial menus, color coding, and explanation tags. Our system encodes object properties and robot limits into object-oriented structures using a vision-language model, allowing explanation generation on the fly and direct manipulation of virtual twins spatially aligned within a simulated environment. We integrate the end-to-end pipeline with a physical robot and showcase diverse use cases ranging from low-level pick-and-place to high-level instructions. Finally, we evaluate X-OOHRI through a user study and find that participants effectively issue object-oriented commands, develop accurate mental models of robot limitations, and engage in mixed-initiative resolution.