Research
Theory of Mind for Explainable Human-Robot Interaction
Overview Research area: Human–robot interaction (HRI), Theory of Mind (ToM), and Explainable Artificial Intelligence (XAI). Technical level: Beginner-Friendly. This is a conceptual/position paper with

- arXiv
- 2512.23482
- Published
- 2025-12-29
- Authors
- Marie S. Bauer, Julia Gachot, Matthias Kerzel, Cornelius Weber, Stefan Wermter
AI summary
Overview
Research area: Human–robot interaction (HRI), Theory of Mind (ToM), and Explainable Artificial Intelligence (XAI).
Technical level: Beginner-Friendly. This is a conceptual/position paper with no new experiments, datasets, or model architectures — its core artifact is a qualitative evaluation of prior ToM-in-HRI studies against an existing XAI evaluation framework.
Scope: The paper argues that Theory of Mind as used in robotics should be treated and assessed as a form of XAI, and it tests that claim by scoring eight recent ToM-in-HRI studies against the eValuation XAI (VXAI) framework and its seven desiderata.
What This Paper Is About
Robots increasingly need to explain their behavior to the people around them, and Theory of Mind — the ability to attribute beliefs, desires, and intentions to others — is often proposed as the mechanism that lets robots do this in a user-friendly way. The problem the authors identify is that ToM studies in HRI claim benefits such as improved understanding, trust, and collaboration, but almost never check whether the explanations a robot gives actually match the robot's real internal reasoning. The goal of the paper is to show that ToM should be evaluated with the same rigor as XAI, and to propose embedding ToM inside an XAI framework so that explanations are both faithful to the system and understandable to users.
Key Contributions
-
Reframing ToM in HRI as a form of XAI. The authors argue that because both fields aim to make internal reasoning understandable to humans, ToM should be evaluated using XAI criteria rather than only user-satisfaction measures.
-
A systematic assessment of eight ToM-in-HRI studies against the VXAI framework. Using the seven VXAI desiderata (Parsimony, Plausibility, Coverage, Fidelity, Continuity, Consistency, Efficiency), the paper builds Table 1 to show which criteria each study addressed.
-
Identification of a specific gap around explanation fidelity. The paper reports that no study examined the model's internal reasoning process, leaving it unclear whether the explanations given to users accurately reflect the model's behavior — a risk of misleading users.
-
A proposed integration direction. The authors propose combining ToM's user-centered perspective with model-centered XAI techniques, and point to possible mechanisms such as behavior trees and explainable reinforcement learning (XRL), building on existing ToM integrations that use Bayesian reinforcement learning to model user behavior.
Main Findings
-
Parsimony and Plausibility are universally met: All of the ToM studies evaluated satisfy these two desiderata, meaning they conducted user-centered experiments and assessed whether their explanations were perceived as believable.
-
Continuity and Consistency are rarely met: Only two studies meet both criteria. The paper attributes this to experiments either not reporting the number of participants or involving fewer than 100 participants, which limits scaling to real-world applications and hurts reproducibility (the VXAI mapping rule is that a minimum of 100 human participants is required for these two criteria to count as evaluated).
-
Coverage is never met: None of the studies reported the number of successful versus unsuccessful interactions, so the Coverage desideratum is unaddressed across the board.
-
Fidelity is never met: None of the studies examined the internal reasoning process of the model. The authors flag this as the central problem, since without it, whether explanations reflect actual model behavior remains unclear.
-
Efficiency is partially addressed: The paper's text does not state an explicit count for the Efficiency desideratum, although Table 1 marks it for several of the studies; the VXAI mapping rule counts Efficiency as evaluated when computational implementation details are provided.
-
Humans can read robot behavior, but only within expectations: One study found that humans interpret robot behavior similarly to human behavior when robots display distinct and interpretable social cues, but this understanding diminishes when a robot's cues deviate from human expectations (Banks 2020).
-
LLMs are not reliable ToM agents: One study showed that while large language models can be a useful tool in human–robot interaction (Becker et al. 2025), they do not function as reliable ToM agents (Verma et al. 2024), supporting the case for explicit ToM-like mechanisms in robotic systems.
-
ToM-equipped robots are perceived better, but explanations are uneven: Robots with ToM capabilities were perceived more positively (Mou et al. 2020), especially when assistance aligned with users' goals (Cantucci and Falcone 2022); robots that reason about human beliefs were seen as more helpful and socially competent (Shvo et al. 2022) and more trustworthy (Angelopoulos et al. 2025). However, when providing explanations, robots may fail to enhance user understanding or improve decision-making, because not all explanations are equally effective (Yuan et al. 2022), whereas approaches implementing multiple levels of explanation improved user comprehension and the interaction (Kerzel et al. 2022).
-
None of these studies used XAI-specific criteria: Although several studies evaluate human–AI collaboration and occasionally describe their work as XAI, none assess it with XAI criteria, and none explicitly integrate ToM with XAI.
Methodology in Plain English
The authors did not run a new robot experiment. Instead, they took an existing, published checklist for judging explainable AI — the eValuation XAI (VXAI) framework by Dembinsky et al. (2025), which consolidates the main evaluation criteria from recent reviews — and used its seven desiderata as scoring criteria for prior ToM work in HRI.
They then went through recent ToM-in-HRI studies and asked, for each one, whether the paper explicitly addressed each desideratum. To keep this consistent, they used fixed mapping rules: a human evaluation counts as addressing Parsimony and Plausibility; reporting successful versus failed interactions counts as addressing Coverage; examining the model's internal reasoning counts as addressing Fidelity; at least 100 human participants counts as addressing both Continuity and Consistency; and providing computational implementation details counts as addressing Efficiency. The results were collected into a single table showing which criteria each study met, and the unmet criteria were used to motivate the proposed integration of ToM into XAI. Definitions of all seven desiderata are provided in the paper's Appendix A.
Why This Matters
Impact on research. The paper draws a line between claims and evidence in ToM-for-HRI: if a system's purpose is to explain, it should be judged by explanation criteria, especially fidelity — whether what the user is told matches what the system actually did. It also pushes back on the dominant direction of XAI research, which the authors say focuses predominantly on the AI system itself and often lacks user-centered explanations.
Real-world applications (contexts implied by the paper's discussion of HRI):
- Collaborative or assistive robots that must explain why they took an action, so users can judge whether to trust and follow it.
- Social robots whose behavior needs to align with typical human expectations, since understanding breaks down when cues deviate from those expectations.
- Robots that reason about a user's beliefs and goals, such as when providing assistance aligned with what the user is trying to achieve.
- Systems using multi-level explanations, which the cited work found improved user comprehension and the interaction.
Industry relevance. Any deployed robot that has to justify its decisions to a non-expert — in care, service, education, or shared industrial workspaces — faces the fidelity problem this paper highlights. The VXAI desiderata also give developers a practical reporting checklist (participant numbers, success/failure counts, implementation details) that maps directly onto documentation and compliance-style evidence.
Future Directions
- Operationalize fidelity for ToM. The central open question is how to actually measure whether a robot's explanation mirrors its internal reasoning; the paper positions fidelity as the criterion to make central.
- Concrete integration designs. The authors propose combining ToM's user focus with model-centered XAI techniques, and suggest exploring behavior trees or explainable reinforcement learning (XRL) inside ToM-based systems to support adaptive reasoning while improving explanation fidelity.
- Build on existing modeling approaches. Existing ToM integrations that use Bayesian reinforcement learning to model user behavior are a starting point the authors say future work could extend.
- Raise evaluation standards in ToM studies. Coverage and Fidelity were unaddressed across all reviewed studies and Continuity/Consistency in all but two, leaving room for future work to report successful versus failed interactions, examine internal reasoning, and use larger participant pools.
Target Audience
Researchers and practitioners working at the intersection of human–robot interaction, social robotics, and explainable AI — particularly those designing or evaluating robot explanation mechanisms. It is also useful for XAI researchers who want a user-centered counterweight to model-centered evaluation, and for students or newcomers who need a compact, jargon-light entry point into how ToM is currently applied and assessed in robotics. Because the paper contains no new experiments or quantitative benchmarks, readers looking for datasets, model names, or performance numbers will not find them here.
Authors’ abstract
Within the context of human-robot interaction (HRI), Theory of Mind (ToM) is intended to serve as a user-friendly backend to the interface of robotic systems, enabling robots to infer and respond to human mental states. When integrated into robots, ToM allows them to adapt their internal models to users' behaviors, enhancing the interpretability and predictability of their actions. Similarly, Explainable Artificial Intelligence (XAI) aims to make AI systems transparent and interpretable, allowing humans to understand and interact with them effectively. Since ToM in HRI serves related purposes, we propose to consider ToM as a form of XAI and evaluate it through the eValuation XAI (VXAI) framework and its seven desiderata. This paper identifies a critical gap in the application of ToM within HRI, as existing methods rarely assess the extent to which explanations correspond to the robot's actual internal reasoning. To address this limitation, we propose to integrate ToM within XAI frameworks. By embedding ToM principles inside XAI, we argue for a shift in perspective, as current XAI research focuses predominantly on the AI system itself and often lacks user-centered explanations. Incorporating ToM would enable a change in focus, prioritizing the user's informational needs and perspective.