Research
HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs
Overview Research area: Socially Assistive Robotics (SAR) and Human-Robot Interaction (HRI), combining multimodal perception with large language models for personalized, multi-user dialogue. Technical

- arXiv
- 2601.19839
- Published
- 2026-01-27
- Authors
- Jeanne Malécot, Hamed Rahimi, Jeanne Cattoni, Marie Samson, Mouad Abrini, Mahdi Khoramshahi, Maribel Pino, Mohamed Chetouani
AI summary
Overview
Research area: Socially Assistive Robotics (SAR) and Human-Robot Interaction (HRI), combining multimodal perception with large language models for personalized, multi-user dialogue.
Technical level: Advanced. The paper assumes familiarity with LLM prompting and inference pipelines, multimodal feature extraction (face recognition, speech-to-text, voice activity detection), LLM-as-a-Judge evaluation, and ablation methodology.
Scope: The paper presents HARMONI, a four-module framework that lets a socially assistive robot identify who is speaking, recall and update speaker-specific profiles over time, and generate personalized responses in multi-user settings, validated on three named datasets plus a nursing home user study.
What This Paper Is About
Robots deployed in shared spaces such as nursing homes must interact with several people at once, remember who said what across sessions, and adapt their responses to each individual — yet most existing HRI systems are built for one-on-one, single-session interaction and lack sustained personalization. HARMONI addresses this gap by coupling large language models with a perception, world-modeling, user-modeling, and generation pipeline that tracks both short-term conversational context and long-term user profiles. The goal is to show that such a framework can robustly identify speakers, keep memory current, and remain ethically and privacy aligned in sensitive care settings.
Key Contributions
- A multimodal personalization strategy that adapts responses to diverse user preferences and behaviors by retrieving only the profile features relevant to the current query, rather than dumping the whole profile into the prompt.
- An online user modeling approach for continuous, adaptive personalization, in which profiles are created when an unrecognized speaker appears and refined as new facts arrive in each turn.
- A real-time multi-user memory management system spanning long-term memory (sustained adaptation across sessions) and short-term memory (context within the present interaction).
- An ethical design methodology and scenario-driven evaluation, explicitly balancing personalization against privacy and safety, validated in realistic nursing home environments with elderly participants and gerontology experts.
Main Findings
-
Perception is accurate and fast. Over 45 video recordings, face detection reached 94.90 accuracy (F1 97.40) at 2.8 ms; speaker recognition reached 89.60 accuracy (F1 88.90) at 1.9 ms; user retrieval reached 97.90 accuracy (F1 93.60) at 3.12 ms. The paper describes these as close to 95% for face detection, 90% for speaker recognition, and close to 98% for user retrieval, all with latency under or near 3 ms.
-
Personalization gains hold across model families. On PersonaFeedback, the proposed framework (retrieving only pertinent profile features via query similarity) consistently outperformed baselines in the specific, domain-oriented mode across all evaluated LLMs. In the general conversation mode, improvements were stronger for larger LLMs, while for smaller LLMs supplying the entire user profile directly worked better.
-
Latency is a trade-off, not an obstacle. The personalization pipeline was slower than direct inference (fewer tokens processed) but lower-delay than the baseline that supplies the full profile. Adding short- and long-term memory improves feature extraction but increases latency.
-
Memory helps differently by model. Incorporating both short- and long-term memory improved session-level semantic similarity for Gemma and LLaMA; Mistral benefited more from long-term memory alone; GPT-4o could extract needed profiling information directly from the query without external memory.
-
Two parallel inferences beat one structured-output inference. Using GPT-4o on the LoCoMo dataset, the two-inference configuration generally produced higher ROUGE F1 scores and better session similarity than the single-inference variant, with LLM-as-a-Judge gains up to three points. The exception was the long-term-memory scenario, where single inference was slightly better. Combining short- and long-term memory gave the best ROUGE and cosine similarity; running the two inferences in parallel kept latency down.
-
Open-source models rival GPT-4o on accuracy but fail more often. On LoCoMo, gemma3:12b scored 7.6 session similarity, 7.9 observation similarity with 17% missed observations; gemma3:27b scored 8.3 and 8.41 with 24% missed; mistral-nemo:12b scored 2.68 and 4.62 with 76.7% missed; llama3:70b scored 1.81 and 1.45 with 90.53% missed; GPT-4o scored 8.8 and 9.16 with 1.96% missed. The body text describes this as "18% missed observations for Gemma-3 12B to nearly 80% for Mistral-Nemo 12B," figures that differ slightly from the table values.
-
Users and experts rated the system highly. Twenty participants (13 women, 7 men) aged 65–91 (M = 78.3, SD = 7.02) each completed a 15-minute interaction. The mean SUS score was 82.4 (SD = 15.8), above the 72/100 acceptability threshold and in the "Good" range (>73), approaching "Excellent" (>85). Two gerontology experts rated the interaction 3.95 and 4.05 on a 5-point scale, an average of 4/5.
-
Qualitative weaknesses were specific. The system could not memorize exact appointment times, often substituting vague adverbs such as "soon." Participants were frustrated by the lack of internet access and restricted temporal knowledge (for example, current weather or updated transportation fares). The system's tendency to congratulate users on the quality of their questions drew both positive and negative reactions. Users and experts valued the system answering "I don't know" rather than inventing information.
-
Summary numbers from the conclusion. Speaker identification around 90%, user profile retrieval around 98%, and personalized response generation of +4/10, all reported as consistently surpassing baselines.
Methodology in Plain English
The researchers built a modular system that splits the problem into four cooperating parts rather than one monolithic model.
Perception takes the video stream, separates audio from images using Voice Activity Detection (FastRTC), finds faces with a YOLOv8 face-detection model, and decides who is talking by correlating facial movement, particularly lip motion, using 68 face landmarks.
User modeling transcribes the speech with Whisper-Turbo, encodes the speaker's face with the INSIGHTFACE buffalo_l model, and compares that encoding against a user database. An unrecognized face creates a new profile; a known face retrieves the existing profile plus relevant past memories found by semantic similarity using the google/EmbeddingGemma-300m text encoder. The speaker's image, transcript, and audio also feed a feature extractor that estimates demographic and affective attributes such as age, gender, and emotion, which initialize or update the profile. A large language model (google/gemma3-27b) reasons over these features.
World modeling keeps short-term conversational state and the set of users present, retrieving the relevant world segment with the same text embeddings used in user modeling, which the authors say reduces hallucination and lets the system signal uncertainty.
Generation takes the query, retrieved memory, profile features, and world state, and produces a response conditioned on all of them. It also filters sensitive or personally identifiable information, enforces safety constraints against harmful or biased output, and respects user-defined privacy preferences.
The system runs with a user interface that displays updated profiles and dialogue history for explainability. Gemma3-27B is served via the Ollama provider; a second inference pass updates the active user's profile from new statements while the first generates the response.
Scenarios were designed to mimic nursing home shared rooms, including a case where a second resident interrupts a reminder meant for another resident and the robot must resolve references such as "the appointment" to the correct person and switch context mid-conversation.
Evaluation was organized around five research questions: (Q1) user detection and profile retrieval on a custom 45-video dataset; (Q2) memory and feature extraction on LoCoMo; (Q3) comparison against prompt-engineered plain LLMs with user models on PersonaFeedback, translated from Chinese to English to French via OPUS; (Q4) latency with open-source and closed-source LLMs; (Q5) real-world nursing home performance and patient opinion. Six models were compared: Gemma3-4B, Gemma3-12B, Mistral-Nemo-12B, LLaMA3-70B, GPT-OSS-20B, and GPT-4o. Mixtral-8x7B (57B parameters, 7B active) served as the LLM-as-Judge. Three personalization settings were tested (Direct Inference, User Modeling, Feature Selection), and an ablation compared single-inference structured output against two sequential inferences. The framework was also integrated with the Mirokai robot from Enchanted Tools through a Python bridge using the pymirokai library and REST/WebSocket APIs, with audio and video captured by the robot sent to HARMONI's interface endpoints and the generated text voiced by the robot's own text-to-speech. Ethical approval for human participants was obtained.
Why This Matters
Impact on research. The paper argues that prior work treats multimodal personalization, multi-user memory management, and ethical safeguards in a disjoint manner, and that HARMONI is a unified approach addressing all three. It moves SAR research beyond single-user, single-session interaction toward systems that handle interruptions, speaker changes, and cross-session memory, while reporting failure rates for open-source models that matter for anyone planning local, privacy-preserving deployments.
Real-world applications.
- Nursing homes and assisted-living facilities where a robot must distinguish among residents sharing a room and keep each person's history separate.
- Therapy centers and group rehabilitation programs where the robot must adapt to several participants in one session.
- Hospital or care settings where reminders about appointments and personal details must stay private and be delivered to the right person.
- Education and collaborative work environments, which the authors name as broader domains with potential.
Industry relevance. The framework is deployable on commodity components (face detection, Whisper-Turbo, embedding models, and open-weight LLMs served through Ollama) and has been demonstrated on a commercial social robot. That makes it directly relevant to robotics companies building care or service robots, and to teams weighing open-source models against closed APIs — the paper quantifies both the accuracy parity and the higher task-failure rates of the open models.
Future Directions
- Extending to multi-agent interactions, moving beyond human users to settings with multiple robotic agents.
- Incorporating richer affective and contextual signals into the personalization pipeline.
- Studying the long-term impact of personalization on trust and engagement, rather than measuring satisfaction in a single 15-minute interaction.
- Improving the deployed implementation, specifically enabling real-time audio and video streaming instead of short recordings, avoiding stored data for privacy, and activating perception modules only when user activity is detected.
Target Audience
Researchers and practitioners in socially assistive robotics and human-robot interaction who are designing systems for multi-user, long-term deployment; engineers building LLM-based dialogue pipelines who need concrete evidence on memory configuration, single-versus-multiple inference design, and latency trade-offs; and teams in healthcare robotics evaluating whether open-source models can substitute for closed APIs in privacy-sensitive care settings. Clinicians, gerontology experts, and ethics reviewers involved in deploying assistive robots around older adults will also find the user study and the qualitative failure analysis directly useful.
Authors’ abstract
Existing human-robot interaction systems often lack mechanisms for sustained personalization and dynamic adaptation in multi-user environments, limiting their effectiveness in real-world deployments. We present HARMONI, a multimodal personalization framework that leverages large language models to enable socially assistive robots to manage long-term multi-user interactions. The framework integrates four key modules: (i) a perception module that identifies active speakers and extracts multimodal input; (ii) a world modeling module that maintains representations of the environment and short-term conversational context; (iii) a user modeling module that updates long-term speaker-specific profiles; and (iv) a generation module that produces contextually grounded and ethically informed responses. Through extensive evaluation and ablation studies on four datasets, as well as a real-world scenario-driven user-study in a nursing home environment, we demonstrate that HARMONI supports robust speaker identification, online memory updating, and ethically aligned personalization, outperforming baseline LLM-driven approaches in user modeling accuracy, personalization quality, and user satisfaction.