Research
Explicit World Models for Reliable Human-Robot Collaboration
Explicit World Models for Reliable Human-Robot Collaboration Overview Research area: Human-robot collaboration (HRC), embodied AI, and neuro-symbolic AI — specifically the reliability of embodied agen

- arXiv
- 2601.01705
- Published
- 2026-01-05
- Authors
- Kenneth Kwok, Basura Fernando, Qianli Xu, Vigneshwaran Subbaraju, Dongkyu Choi, Boon Kiat Quek
AI summary
Explicit World Models for Reliable Human-Robot CollaborationOverview
Research area: Human-robot collaboration (HRC), embodied AI, and neuro-symbolic AI — specifically the reliability of embodied agents operating under sensing noise, ambiguous instructions, and dynamic human-robot interaction.
Technical level: Intermediate. The paper is a position paper rather than an experimental study; it is conceptual and literature-driven, but it presumes familiarity with terms such as common ground, perceptual grounding, neuro-symbolic architectures, and world models.
Scope: The paper argues that reliable human-robot collaboration should be built on an explicit, accessible, continuously updated world model that serves as common ground between humans and robots, rather than on end-to-end black-box models or formal verification alone.
What This Paper Is About
The authors take the position that reliability in embodied AI cannot be treated as a fixed property of a learned model, because human environments are inherently social, multimodal, and fluid, and because reliability only has meaning relative to the goals and expectations of the humans involved. Instead of pursuing formal verification for model predictability and robustness, they argue that robots should construct and maintain an explicit world model that captures shared understanding of tasks, communications, and environments. The goal is to align robot behaviour with human expectations by making the robot's interpretation of environmental state and human intention explicit, inspectable, and updatable in real time.
Key Contributions
-
A reframing of reliability for embodied AI. The paper redefines reliability as contextually determined and relational — meaningful only with respect to the goals and expectations of the humans in the interaction — in contrast to approaches that place the burden of reliability on learning verifiably correct models for each task and situation.
-
A human-inspired account of common ground for HRC. It reviews how humans build shared understanding through continuous integration of multimodal cues (gaze, gestures, prosody, movement dynamics, and contextual knowledge), and traces this concept from language and cognition studies into human-AI teaming constructs such as shared mental models, knowledge representation, schema, and situation awareness.
-
A review of perceptual grounding and multimodal interaction research. The paper surveys visual grounding work and the authors' own prior contributions — the Task-oriented Collaborative Question Answering (TCQA) benchmark, M2GESTIC, COSM2IC, and Ges3ViG — showing that reliability emerges from dynamic, context-dependent coordination of modalities rather than from rigid pipelines, and that Large Language and Multimodal Models remain limited by their disembodiment from the physical world.
-
A review of explicit world modelling in AI, culminating in a call to action. It contrasts handcrafted symbolic cognitive architectures with emerging neuro-symbolic approaches such as NeSyC, Knowledge Module Learning (KML), and the PKR-QA benchmark, and concludes with a call for multidisciplinary work to build lightweight yet representative explicit world models.
Main Findings
-
Reliability is not a property of a model alone: The paper argues that in social, multimodal, and fluid human environments, reliability is contextually determined and only meaningful relative to the goals and expectations of the humans involved in the interaction.
-
Isolated perception is insufficient for collaboration: Foundational social robotics work cited in the paper shows that joint attention enables robots to interpret human referential cues, that non-verbal behaviours improve efficiency and robustness in human-robot teamwork, and that robots must act expressively, not just efficiently, producing legible motion that communicates intent.
-
Continuous monitoring of both deliberate and involuntary cues matters: The paper cites work showing that robots need to continuously monitor human behaviour expressed both through conscious actions and language and through unconscious or involuntary non-verbal cues in order to actively infer human intentions.
-
Multimodal grounding reduces ambiguity: The authors' own prior results are reviewed — a distance-weighted understanding of pointing gestures significantly reduces ambiguity in comprehending natural multimodal human instructions (M2GESTIC), eye gaze provides strong cues for predicting referents and action steps during joint tasks, and adaptive real-time multimodal fusion that prioritises gesture or linguistic structure depending on context shows reliability emerging from dynamic coordination (COSM2IC).
-
Task-oriented benchmarks were proposed because VQA is inadequate: The paper states that Visual Question Answering is inadequate at capturing the dynamic, multimodal, and task-specific context of HRC, motivating the Task-oriented Collaborative Question Answering (TCQA) benchmark for quantitative evaluation of grounding methods in HRC tasks.
-
Symbolic world models do not scale; purely learned ones are disembodied: The paper reports that most symbolic approaches depend heavily on human handcrafting and cannot scale to the complexity of real worlds, while LLM/LMM-based approaches for affordance reasoning, coordination, and human goal reasoning face challenges owing to their intrinsic disembodiment from the physical world.
-
Neuro-symbolic world models are emerging as the alternative: NeSyC combines LLM generative creativity with symbolic solver precision in a feedback loop between inductive inference (via LLMs) and deductive validation (via Answer Set Programming), while KML and PKR-QA encode procedural knowledge in a knowledge graph linking tasks, steps, actions, objects, tools, and purposes, yielding interpretable, verifiable stepwise reasoning traces.
-
No quantitative results are reported: This is a position paper that reviews related work; it reports no new experiments, datasets, benchmark scores, or numeric performance figures.
Methodology in Plain English
This is a position paper, so the approach is argument and literature synthesis rather than experimentation. The authors first lay out the problem — robustness under sensing noise, ambiguous instructions, and human-robot interaction — and then reject the assumption that reliability can be achieved purely by verifying that learned models behave predictably. They build their case in two literary strands. The first strand draws on human cognition and social robotics to explain how people establish common ground: by continuously integrating gaze, gesture, prosody, movement, and contextual knowledge, by sharing attention, and by communicating intent through expressive action. The second strand reviews how AI systems have tried to represent the world explicitly — starting with symbolic cognitive architectures that manipulate formally defined symbols, then moving to newer neuro-symbolic systems that learn their own internal representations from sensorimotor experience while retaining interpretable, structured reasoning. Along the way, the authors cite their own prior work on multimodal referential grounding and procedural knowledge learning as evidence for their claims. They then combine the two strands into a position: that the common ground built through interaction should be materialised as an explicit world model, which resolves ambiguity and subjectivity by explicitly committing to particular interpretations of environmental states and human intentions. They close with requirements and a call for multidisciplinary collaboration.
Why This Matters
Impact on research: The paper pushes back against the dominant end-to-end, black-box trend in embodied AI and against treating reliability as a formal verification problem alone. It proposes an alternative research programme organised around explicit, interpretable, self-evolving world models, and it connects previously separate literatures — social robotics and common ground research on one side, neuro-symbolic cognitive architectures on the other. It also identifies a concrete design tension for the field: such models must be light-weight enough for real-time updating yet representative enough to capture rich social, multimodal, and fluid interactions.
Real-world applications (as implied by the paper's framing):
- Collaborative robots in shared workspaces, where legible motion and shared attention improve predictability and coordination with human co-workers.
- Task-oriented human-robot teamwork in which robots must interpret multimodal instructions that combine speech with pointing gestures and gaze.
- Assistive and service robotics operating in human environments where instructions are ambiguous and intentions must be inferred from both deliberate and involuntary cues.
- Industrial and manufacturing settings represented by the authors' institutional collaborations, where robots must coordinate with human operators on procedural tasks.
Industry relevance: Two of the three author affiliations are corporate and national research organisations (A*STAR institutes in Singapore and a Samsung Electronics mechatronics research team in Korea), signalling industrial interest in deploying collaborative robots that behave comprehensibly and predictably around people. The paper's argument that reliability must be maintained on the fly, rather than certified once per task, speaks directly to deployment contexts where tasks and environments change faster than models can be re-verified.
Future Directions
-
Building world models that are both lightweight and representative. The paper explicitly names the tension between real-time updatability and the richness needed to capture social, multimodal, and fluid human-robot interaction as an open requirement rather than a solved problem.
-
Realising the shift from opaque models to explicit world models. The core call to action is to move from black-box end-to-end control toward explicit world models that provide common ground for guiding reliable collaborative behaviour in human-robot teams.
-
Multidisciplinary collaboration. The authors state that the challenges they identify will require contributions from the diverse communities present in this Bridge, implying that no single discipline — robotics, cognitive science, NLP, or symbolic reasoning — can resolve them alone.
-
Scaling explicit world modelling beyond handcrafting. The paper points to neuro-symbolic continual learners such as NeSyC and knowledge-module approaches such as KML as a trajectory for developing interpretable, self-evolving world models, raising the question of how far that trajectory can extend toward open-domain, real-world deployment.
Target Audience
This paper is best suited to robotics and embodied AI researchers, particularly those working on human-robot collaboration, human-robot teaming, and multimodal grounding. It will also interest neuro-symbolic AI and cognitive architecture researchers, since it positions explicit world models as the bridge between symbolic reasoning and learned perception. Human factors and HRI researchers studying common ground, shared mental models, and situation awareness will find the framing relevant, as will industry practitioners and research programme managers evaluating whether end-to-end learning is a viable route to deployable collaborative robots. Because it is a position paper without experimental results, readers seeking benchmark numbers or implementation details will need to follow the cited prior work instead.
Authors’ abstract
This paper addresses the topic of robustness under sensing noise, ambiguous instructions, and human-robot interaction. We take a radically different tack to the issue of reliable embodied AI: instead of focusing on formal verification methods aimed at achieving model predictability and robustness, we emphasise the dynamic, ambiguous and subjective nature of human-robot interactions that requires embodied AI systems to perceive, interpret, and respond to human intentions in a manner that is consistent, comprehensible and aligned with human expectations. We argue that when embodied agents operate in human environments that are inherently social, multimodal, and fluid, reliability is contextually determined and only has meaning in relation to the goals and expectations of humans involved in the interaction. This calls for a fundamentally different approach to achieving reliable embodied AI that is centred on building and updating an accessible "explicit world model" representing the common ground between human and AI, that is used to align robot behaviours with human expectations.