Research
SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented Reality
Overview Research area: Human-Computer Interaction (HCI) — specifically wearable augmented reality assistants, speech interaction, spatial memory, and LLM-driven intent inference. Technical level: Int

- arXiv
- 2602.00793
- Published
- 2026-01-31
- Authors
- Yoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang, Arie Kaufman
AI summary
Overview
Research area: Human-Computer Interaction (HCI) — specifically wearable augmented reality assistants, speech interaction, spatial memory, and LLM-driven intent inference.
Technical level: Intermediate. The paper is a system-and-user-study contribution: the technical pipeline is described at an architectural level (agents, retrieval, fusion, confidence scoring) rather than with deep math, while the evaluation leans on standard HCI statistics (repeated-measures ANOVA, Friedman, Wilcoxon signed-rank).
Scope (one sentence): The paper designs and evaluates SpeechLess, a wearable AR assistant that lets users "speak less" — from full sentences down to micro- or zero-utterance — by extrapolating missing intent from a personalized spatial memory bound to space, time, activity, and referents.
What This Paper Is About
Speaking aloud to a wearable AR assistant in public is socially awkward, and repeating the same request every day is unnecessary effort. The authors ask whether a system can let users deliberately under-specify a query — or say nothing at all — and still get the right answer, by remembering how, where, when, and about what the user previously asked. SpeechLess is their proof-of-concept answer: speech becomes a granularity control knob rather than an all-or-nothing channel, backed by long-term personal spatial memories that fill in the missing pieces of intent.
Key Contributions
- An interaction paradigm for everyday wearable AR: users control query intent granularity through speech, adapting to social and situational context, rather than being forced into either full articulation or purely system-triggered proactivity.
- A personal spatial memory representation: a spatially grounded memory design that binds digital information with personal multimodal context — speech, space, time, actions, and referents — in the real world to enable context-aware intent resolution.
- Empirical findings on current smart wearables: a longitudinal in-the-wild formative study with Meta RayBan glasses exposing practical limitations of speech-based wearable AI in public and private spaces.
- Demonstration and validation: evidence that convenience and social acceptability improve under controlled lab and socially constrained ("crowds") conditions, without substantially degrading perceived usability or intent resolution accuracy.
Main Findings
- Formative study — social norms override need (7-day study, N=8, Meta RayBan glasses): Participants avoided speaking aloud even when they wanted assistance, with verbal interaction perceived as inappropriate in libraries, classrooms during lecture, and office settings (theme reported as N=9 in the paper, though the study recruited N=8). One participant: "I wanted to look something up during class, but I just couldn't without disturbing everyone." Workplace-norm hesitation was also reported (N=3).
- Formative study — articulation burden: Repetitive wake-word invocation ("Hey Meta") was described as cumbersome (N=6); several participants switched to a physical tap on the side of the glasses. Reluctance to speak every question aloud was reported by N=5.
- Formative study — few novel use cases: Daily routines left little room for new interactions (N=6), with repetitive work-to-home commutes yielding mostly weather, traffic, and bus-schedule queries; N=2 noted a lack of use cases when a smartphone could do the same task.
- Formative study — social awkwardness and privacy: Blinking indicator lights paradoxically drew more attention; bystanders misjudged participants as recording (N=5), one participant was assumed to be seeking human assistance, and the uncommon form factor plus verbal queries drew unwanted attention (N=6).
- Formative study — hardware constraints: Battery lasting less than 30 minutes during consecutive video recording (N=1), paired smartphone overheating during data synchronization (N=2), and limits in recording resolution, connection reliability, and form-factor convenience (N=4).
- Study A — perceived relevance (18 participants, ages 21–25, μ=23.0, σ=1.19; Meta Quest 3 with passthrough AR): No significance across modes; means were Full 5.94 (±0.87), Partial 5.44 (±1.04), Zero 5.11 (±1.68) on a 7-point scale. Overall spatial adaptability was rated 5.78 (±1.11), significantly above the neutral score of 4 (Wilcoxon, p<.001).
- Study A — Task 1 (Remembrance) cognitive load: Mental demand ranked Full 47.78 (±24.87), Zero 29.44 (±27.96), Partial 25.00 (±18.86) (Friedman χ²=23.41, p<.001; Full>Partial p<.001, Full>Zero p<.01). Physical demand: Full 42.78 (±22.18), Zero 28.33 (±23.33), Partial 22.22 (±16.29) (χ²=19.00, p<.001). Effort: Full 40.00 (±20.86), Zero 27.22 (±21.91), Partial 21.11 (±11.83) (χ²=20.22, p<.001), though only Full>Partial (p<.001) was significant pairwise.
- Study A — Task 2 (Comparison) trade-off: Partial lowered most loads but increased temporal demand — Mental demand Full 37.22 (±16.74) vs Partial 19.44 (±12.59), p<.001; Physical demand Full 34.44 (±10.42) vs Partial 18.33 (±9.24), p<.001; Effort Full 32.22 (±20.74) vs Partial 16.67 (±6.86), p=.006; but Temporal demand was higher for Partial 34.44 (±23.32) than Full 18.89 (±12.31), p=.025. One participant (P10) attributed this to needing additional retrieval attempts.
- Study A — usability: Zero was rated the most difficult to use (5.56 ±1.95), followed by Full (5.28 ±1.74) and Partial (4.78 ±1.59), but differences were not statistically significant. Mean perceived usability across all conditions was "Good" with a SUS score of 75.42. Usefulness for reducing articulation burden scored 5.39 (±1.09) for Partial and 5.22 (±1.48) for Zero.
- System behavior: The Memory Retriever surfaces up to five candidates (k=5) using GPS proximity, semantic similarity, and temporal recency, combined via Reciprocal Rank Fusion; the Answer Composer produces responses capped at up to 30 words, validated with a 1–10 confidence score, with low-confidence responses routed to a Yes/No user verification interface.
- Not reported in the provided content: Study B (the in-the-wild longitudinal deployment) participant counts, duration, and quantitative results are not included in the text available here; it is described as examining practical adoption, social acceptability, and privacy, and its in-the-wild usages inspired the six illustrated use-case scenarios.
Methodology in Plain English
The authors began with a week-long field study: eight people wore Meta RayBan glasses for at least six hours a day (not consecutively) in both private and public settings, logged their daily experiences, and were interviewed at the end; three co-authors independently coded the feedback into themes via inductive thematic analysis.
Those themes drove the design goals: decouple assistance from mandatory speech, bind routine queries to personal context, keep the user in the loop for uncertain inferences, and use a minimal on-demand interaction model rather than continuous sensing — explicitly no SLAM and no gesture recognizer requiring continuous frame access. As a fallback, four button triggers mirror the Meta RayBan tap gestures (single, double, triple, tap-and-hold).
The system then runs as a six-module pipeline. A Query Decoding Agent classifies each input as Question Answering (subtyped as Full, Partial, or Zero Utterance), Remembrance, or Removal. A Contextual Dimension Encoder builds a "dimension sketch" of space label, scene description, referent/object-of-focus, temporal moment, action intent, and raw speech transcription. The Memory Store Manager writes these sketches plus query and response as episodic "spatial memories." For Partial and Zero queries, the Memory Retriever compares the current sketch against stored episodes and returns up to five candidates, updating live content (such as transit or weather) through the Google Custom Search API when needed, or falling back to the LLM's general knowledge when no prior memory exists. The Answer Composer aggregates the candidates with the current view to generate a short, context-adaptive answer with a rationale, and low-confidence answers are confirmed by the user.
Study A isolated the interaction paradigm from device ergonomics by using a controlled HMD platform. Eighteen participants used three utterance modes (Full as baseline, Partial, Zero) in a within-subjects, counterbalanced design across three physical mock-up environments (Office, Break Room, Bistro), with three personas and 45 pre-seeded personal memories each. They completed two tasks — T1 Remembrance across three physically different locations, and T2 Comparison across spaces or objects (Zero was excluded from T2 for lack of cues). Each participant chose 18 or more spatial memories (3 locations × 3+ trials × 2 task types) and completed RTLX (0–100), 7-point Likert, SUS, and semi-structured interviews. Responses were analyzed with repeated-measures ANOVA where normality held (Shapiro-Wilk), otherwise Friedman with Bonferroni-corrected pairwise tests or Wilcoxon signed-rank. The study was IRB-approved (1173920), and participants were compensated $15 for a 90-minute session.
Why This Matters
Research impact. The paper reframes speech in wearable AR as a variable-granularity control rather than a fixed input modality, and positions personal spatial memory as more than a reminder store — it becomes the mechanism that fills implicit intent gaps in under-specified queries. It also supplies longitudinal in-the-wild evidence about why current speech-based wearables fail socially, which is directly usable by designers of next-generation assistants.
Real-world applications (as demonstrated in the paper's use cases):
- Home routines: recalling "plant?" on a Tuesday instead of re-asking for a plant-watering reminder.
- Commute: retrieving an up-to-date bus schedule ("M11 bus?") without speaking loudly in a crowd.
- Classroom or library: quietly flagging "Assignment" while looking at course material, then recalling it later on the same material.
- Grocery and bistro contexts: recalling "unsweetened soy milk from Walmart" across the home-to-store spatial gap, or asking "sugar?" about a different sauce while the system adapts the prior sugar-content query.
- Maintenance: recalling a month-old note such as connecting a wire to the port second from the left with a single glanced utterance.
Industry relevance. The findings bear directly on smart-glasses makers (the paper builds on Meta RayBan Gen-1 constraints and Meta Quest 3 for evaluation), LLM assistant vendors implementing retrieval-augmented personalization, and privacy/AR stakeholders concerned with bystander exposure and unintended disclosure. The hardware findings — sub-30-minute battery life under consecutive video recording, overheating during synchronization, and the social cost of wake words — set concrete targets for the next device generation.
Future Directions
- Strengthen the advanced-comparison case. Partial utterance increased temporal demand in T2 because retrieval sometimes required extra attempts; improving retrieval when detailed keywords are absent is an explicit open problem, as is extending Zero utterance into comparison tasks where it was excluded.
- Scale the in-the-wild evidence. The provided text does not report the results, sample size, or duration of Study B; a fuller longitudinal deployment is the natural next step for validating social acceptability and privacy claims outside the lab.
- Reduce dependence on continuous context acquisition. The design deliberately avoids SLAM and continuous gesture recognition to conserve battery; finding ways to enrich spatial memory under the same practical constraints remains open.
- Widen memory robustness and management. Because storage is gated by user verification and low-confidence answers trigger correction rather than silent recording, questions remain about how well the memory store scales, how users curate it over long periods, and how "fresh knowledge" fallbacks with no prior memory should be validated.
Target Audience
Researchers and practitioners in HCI, XR/AR, and wearable computing; interaction designers building speech and multimodal assistants; engineers working on LLM-based retrieval, RAG pipelines, and personal memory systems; and privacy or social-acceptability researchers studying always-on wearable devices. The paper is most valuable to readers interested in how interaction design and context-aware memory can be combined to make public, low-effort AR assistance practical.
Authors’ abstract
Speaking aloud to a wearable AR assistant in public can be socially awkward, and re-articulating the same requests every day creates unnecessary effort. We present SpeechLess, a wearable AR assistant that introduces a speech-based intent granularity control paradigm grounded in personalized spatial memory. SpeechLess helps users "speak less," while still obtaining the information they need, and supports gradual explicitation of intent when more complex expression is required. SpeechLess binds prior interactions to multimodal personal context-space, time, activity, and referents-to form spatial memories, and leverages them to extrapolate missing intent dimensions from under-specified user queries. This enables users to dynamically adjust how explicitly they express their informational needs, from full-utterance to micro/zero-utterance interaction. We motivate our design through a week-long formative study using a commercial smart glasses platform, revealing discomfort with public voice use, frustration with repetitive speech, and hardware constraints. Building on these insights, we design SpeechLess, and evaluate it through controlled lab and in-the-wild studies. Our results indicate that regulated speech-based interaction, can improve everyday information access, reduce articulation effort, and support socially acceptable use without substantially degrading perceived usability or intent resolution accuracy across diverse everyday environments.