Research
From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
Overview Research area: Human-Computer Interaction / Extended Reality — specifically remote collaboration, spatial language grounding, Augmented Reality (AR) interfaces, and Large Language Model (LLM)

- arXiv
- 2602.03059
- Published
- 2026-02-03
- Authors
- Yoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie Kaufman
AI summary
Overview
Research area: Human-Computer Interaction / Extended Reality — specifically remote collaboration, spatial language grounding, Augmented Reality (AR) interfaces, and Large Language Model (LLM) reasoning.
Technical level: Intermediate. Readers will get the most from this paper with basic familiarity with AR/MR head-mounted displays, scene-graph or relational-graph representations, and how LLMs are used for parsing and reasoning. No deep mathematics is required.
Scope (one sentence): The paper presents Speech-to-Spatial, a framework that turns speech-only remote-assistance instructions into spatially grounded AR visual guidance by disambiguating the referent through an object-centric relational graph and LLM reasoning, evaluated against a voice-only baseline.
What This Paper Is About
When a remote expert guides a local worker by voice, spatial language such as "this one" or "over there" is under-specified, causing repeated back-and-forth clarification. Prior systems resolve this ambiguity using extra cues such as gesture or gaze, or by requiring a human to draw annotations manually. Speech-to-Spatial instead infers the intended target from spoken references alone and automatically places a persistent AR indicator on that object in the live shared view.
Key Contributions
- An end-to-end pipeline for speech disambiguation. The system resolves the intent behind speech-only instructions using graph-traversal reasoning and provides situated visual guidance through an AR indicator.
- Spatial description patterns in remote instructions. The authors derive five recurring language patterns in spatial descriptions — Direct-feature, Relational, Memory, Chained, and Deictic — grounded in established studies of language and in their own formative study.
- Demonstration on collaborative scenarios. They showcase Speech-to-Spatial for Remote Maintenance, Indoor Navigation, and Personal Assistance.
- Empirical evaluation. They assess and compare Speech-to-Spatial against an existing voice-only remote-assistance baseline on cognitive load and usability, with implications for integration into existing methods.
Main Findings
- Speech-only referencing is structured, not arbitrary. A formative study with 9 participants (8 male, 1 female; aged 27–34; P1–P9) using Zoom screen sharing isolated four recurring reference types out of 187 observation notes.
- Direct Feature was the most common pattern. Instructors described the target through intrinsic attributes in 57.6% of observations (e.g., "the red file," "the PDF file," "the file named A"). It works when the distinguishing attribute is conspicuous but is fragile when multiple items share features.
- Relational referencing was the second most common. It accounted for 31.2% of occurrences and uses a conspicuous anchor object as a landmark (e.g., "the one to the left of the yellow file"). It requires the follower to identify the anchor first; when that failed, instructors fell back to cursor-based micro-guidance such as "move a little more to the left."
- Memory-based referencing appeared at 11.2%. Instructors referred back to objects mentioned or manipulated earlier. One participant reported this pattern caused additional cognitive overhead from recalling interaction history.
- Chained referencing is composite. It layers Direct Feature, Relational, and/or Memory cues in a single utterance (e.g., "the folder behind the one we selected earlier"). Its statistics are reported as broken down into the individual patterns rather than as a separate percentage.
- Deictic references were most prominent with annotation tools but excluded. Phrases like "that one" or "it" depend on an additional cue such as visual marking or gesture, so they were treated as out of scope for a speech-only system.
- Speech-to-Spatial reportedly improves outcomes over a voice-only baseline. The paper states that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to the conventional voice-only baseline. The specific numeric results of the 18-participant evaluation are not included in the available content.
Methodology in Plain English
The researchers first ran a formative study to learn how people actually talk when guiding someone remotely. They paired participants as instructor and follower on a shared 2D desktop view over Zoom, running a 15-minute session with 30 instructions per session. The first 15 tasks used speech only; the remaining 15 allowed an annotation tool, so the team could see whether strategies shifted when visual marking became available. The first author recorded transcripts and notes, then thematically coded them.
Those findings drove the system design. The core idea is an object-centric relational graph: every visible physical object becomes a node that stores its own features (color, class label, shape, 6DoF pose), its scene context, and a memory field recording actions, actors, and timestamps. Nodes are linked by six spatial relations — "left," "right," "above," "below," "in-front-of," and "behind-of" — and two objects are considered in relation only when they are within a half-meter radius (r = 50 cm). Relations are computed from each object's 3D center point, derived from its axis-aligned 3D bounding box.
At runtime, the system transcribes speech and parses the utterance into a fixed structure: a target object, one or more anchor objects, their class/label and description/features, a relational phrase, and any action, intent, or temporal cues. Generic nouns such as "thing" or "it" are not treated as a target label but as a question about the referent. To handle free-form descriptive language, each node's attributes are embedded and compared to the parsed attributes using cosine similarity; the top five candidates (k = 5) are kept, then an LLM reasons over that reduced set to pick the referred node. Object-level frustum and occlusion culling discard off-view or occluded nodes before similarity checks, using a depth test that compares camera-to-object distance against ray-hit distance plus the target object's scale. If reasoning conflicts or fails, an evaluation LLM agent checks the result and, on conflict, the system falls back to showing the raw transcription instead of a pointer, to avoid misleading guidance.
A resolved referent gets a directional arrow anchored above it, persistent until an action is performed on it, plus an LLM-summarized instruction panel; multi-step instructions are shown as an alphabetized list (A, B, C). The client runs Unity AR Foundation for capture, voice recording, pose extraction, and AR anchoring, sending JSON nodes to a custom Python server. Whisper handles transcription, GPT-4.1 handles parsing, reasoning, and resolution, text-embedding-3-small generates attribute embeddings (cached and reused), and Gemini 2.5-Flash localizes, segments, classifies, and visually analyzes the scene.
Why This Matters
Impact on research. The paper argues that referent ambiguity can be reduced through structural analysis of how entities relate spatially and semantically, without the sensing burden of gaze, gesture, or avatar tracking. It positions graphs as an interpretable alternative to flat vector representations, since references resolve through explicit, traceable paths. It also extends the "Put-That-There" line of work by approximating the disambiguating role of embodied cues using speech alone.
Real-world applications.
- Remote technical maintenance — an expert's utterance such as "locate the second fuse from the left, just below the green wire" becomes an automatically placed visual overlay for the on-site technician, without the expert manually annotating the feed.
- Indoor navigation — step-by-step verbal directions combining landmarks and relational anchors are converted into situated AR guidance, replacing the listener's mental map with a visible one.
- Personal AI assistants — a visual indicator (for example, a dot on the object of interest) clarifies which referent a spoken query means when the assistant might otherwise pick the wrong object.
- AI customer service — the authors suggest a pathway from chat-bot style interaction toward visually grounded agents that can guide a user from a single shared snapshot.
Industry relevance. The paper points out that many commercial platforms (naming TeamViewer Assist AR and Microsoft Dynamics 365 Remote Assist/Teams) remain dominated by 2D desktop or mobile interfaces with manual pointers, cursors, or human-authored annotations, and that immersive solutions are slowed by specialized hardware requirements such as motion tracking and HMDs. A speech-only disambiguation layer addresses that adoption barrier, and the client-server split offloads computation from the client device.
Future Directions
- Extending to deictic references. Phrases such as "that one" and "it" were the most prominent pattern when an annotation tool was available but were excluded because they depend on visual marking or explicit gestural pointing. The paper states this is discussed in its future work discussion, which is not included in the available content.
- Scaling and generalizing the graph representation. The current spatial graph uses six fixed relational properties and a half-meter relation radius with 3D center points from axis-aligned bounding boxes, described as a simplification; richer spatial modeling is an open question.
- Robustness of the fallback path. When reasoning conflicts or nuances are lost in summarization, the system defers to raw transcription. Understanding how often this happens and how users behave in those moments is a natural next step, and the feasibility analysis section of the evaluation is truncated in the available content.
- Longer-term memory use. Because object nodes persist across sessions, the paper raises the prospect of using accumulated interaction history for longer-term usability and spatial reasoning, including references to objects from previous days.
Target Audience
Researchers and practitioners in AR/XR, remote collaboration, and human-computer interaction who are interested in grounding natural language in physical space; engineers building remote-assistance or tele-assistance products; and readers working on LLM-driven scene understanding, scene graphs, and spatial language disambiguation. It is also useful for designers who need lightweight, hardware-agnostic alternatives to gaze- or gesture-based interaction.
Authors’ abstract
We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance ("a bit more to the right", "now, stop.") during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speechto-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.