Research
RLSLM: A Hybrid Reinforcement Learning Framework Aligning Rule-Based Social Locomotion Model with Human Social Norms
Overview Research area: Socially-aware navigation for autonomous agents, spanning human–robot interaction, reinforcement learning, and cognitive/psychological modeling of personal space. Technical lev

- arXiv
- 2511.11323
- Published
- 2025-11-14
- Authors
- Yitian Kou, Yihe Gu, Chen Zhou, DanDan Zhu, Shuguang Kuai
AI summary
Overview
Research area: Socially-aware navigation for autonomous agents, spanning human–robot interaction, reinforcement learning, and cognitive/psychological modeling of personal space.
Technical level: Intermediate. The reinforcement learning scaffolding (actor-critic, reward shaping, ablation studies) assumes some familiarity with RL, but the paper's core idea — encoding measured human comfort into a reward function — is explained without deep mathematics.
Scope: The paper proposes and evaluates a single hybrid framework (RLSLM) that fuses a rule-based, experiment-derived social locomotion model into the reward function of a reinforcement learning agent, validated through an immersive VR user study against two rule-based baselines.
What This Paper Is About
Agents that move through spaces shared with people must avoid making those people uncomfortable, but existing approaches force a choice: rule-based methods are transparent and cheap yet rigid and often unnatural, while data-driven methods are flexible yet expensive to train, opaque, and hard to align with human intuition. This paper asks whether the two can be merged — using psychological findings about human comfort as the reward signal that drives reinforcement learning — so that an agent learns socially appropriate paths quickly and in a way researchers can still interpret. The authors test this by having real participants stand inside a VR scene as virtual bystanders and rate how comfortable the agent's path made them feel.
Key Contributions
- A hybrid reinforcement learning framework (RLSLM) that embeds a psychologically grounded social locomotion model directly into the reward function, combining the interpretability and prior knowledge of rule-based methods with the adaptability of data-driven ones. The authors state the framework is potentially generalizable to other scenarios with similarly scarce data.
- A measured improvement in user comfort: RLSLM achieves a mean comfort rating of 4.21/5, significantly outperforming the best rule-based baseline (Δrating = 1.12, Bonferroni corrected post-hoc comparisons, P < 0.001), which the authors describe as establishing a new Pareto frontier in the trade-off between comfort and efficiency.
- An interpretable multi-objective reward design that jointly optimizes mechanical energy, goal progress, and social discomfort, with a single tunable social weight (σ) whose effects on behavior are quantified through sensitivity analysis.
- A reusable VR-based human–agent interaction dataset and evaluation pipeline, built in Unreal Engine 5.4, with the code and UE project open-sourced to support benchmarking and reproducibility.
Main Findings
- RLSLM rated most comfortable overall: A repeated-measures ANOVA found a significant main effect of model type on comfort level (F(2,58) = 219.589, P < 0.001, η²G = 0.525).
- RLSLM beats both baselines: Rating scores for RLSLM paths were significantly higher than COMPANION (P < 0.001) and n-Body (P < 0.001) in both single-human and multi-human scenarios.
- Multi-human advantage is model-dependent: Compared with single-human scenarios, RLSLM (P < 0.001) and n-Body (P = 0.008) received significantly higher scores in multi-human scenarios, whereas COMPANION showed no such advantage (P = 0.251).
- Exclusion of a repetitive scene did not change conclusions: Before excluding one repetitive three-human scenario, the ANOVA gave F(2,58) = 228.112, P < 0.001, η²G = 0.534; after exclusion, F(2,58) = 219.589, P < 0.001, η²G = 0.525. The n-Body multi-human comparison shifted from P = 0.032 (before) to P = 0.008 (after), and COMPANION from P = 0.931 (before) to P = 0.251 (after).
- The social weight σ controls detour behavior: In a sensitivity analysis over σ ∈ {0, 0.5, 1.0, 2.0}, higher σ produced greater lateral deviation as measured by Maximum Lateral Distance (MLD). At σ = 0 the agent followed the shortest path; at σ = 2.0 behavior became overly conservative.
- Heading sensitivity depends on the HRSC component: Across 42 specially designed single-human scenarios, removing the heading-relevant social component (HRSC) caused the agent to pass in front of the human in 23 cases (57.76%), versus only 5 cases for the full model.
- Heading-irrelevant components affect path stability: Measured in 21 single-human scenarios, removing either the heading-irrelevant social component (HISC) or the collision avoidance component (CAC) reduced MLD, indicating less stable and less compliant navigation.
- Training is fast and stable: Learning curves show episode rewards increasing and episode lengths decreasing; the multi-human condition converges more slowly due to interaction complexity but still reaches stable performance within the 10,000-step training budget.
Methodology in Plain English
The researchers started from prior psychological experiments that measured how uncomfortable people feel when someone passes near them, and distilled those measurements into a mathematical "discomfort field." This field is orientation-sensitive and asymmetric: it is stronger in front of a person than behind them, and it also accounts for the body's roughly elliptical cross-section, so comfort drops faster along the narrow axis than the wide one. Individual persons are each assigned such a field, and the total discomfort at any point is the sum of the contributions from all nearby people.
That field becomes part of a reward signal. At every step the agent receives a small penalty for the energy it spends moving, a bonus proportional to how much closer it got to its destination, and a penalty proportional to how much social discomfort it causes. Reaching the goal yields a large positive reward; running out of steps or leaving the arena yields a large negative one. A weight σ controls how much the agent cares about social discomfort relative to efficiency.
The agent itself is a standard actor-critic setup trained with Advantage Actor-Critic (A2C) via Stable-Baselines3. Observations concatenate the agent's own position with the relative positions and orientations of the surrounding people, giving an observation vector of dimension 3n+2 for n individuals. Training used RMSprop with a learning rate of 5 × 10⁻⁴, an MLP with five hidden layers of 64–128–256–128–64 units, and an NVIDIA 3090 GPU under a fixed 10,000-step budget per run. The virtual world is a 15 m × 15 m space where the agent moves in discrete 45 cm steps, and trajectories are smoothed with a Gaussian filter because the raw stepwise paths looked jagged.
The evaluation put people inside the loop. Thirty university students and staff (11 male, 19 female, aged 18–29) viewed scenarios through an HTC Vive Pro headset (2,880×1,600 pixels binocular resolution, 90 Hz refresh rate, 110° field of view), standing at the position and orientation of a randomly chosen virtual bystander while a virtual agent walked past. Participants rated comfort on a 1–5 scale after each of 150 trials, drawn from 50 scenarios and three algorithms, with score values and trajectory display order randomized.
Why This Matters
Research impact. The work offers a concrete recipe for coupling cognitive science with machine learning: rather than hand-tuning social rules by intuition or letting a network infer them from data, it imports parameters fitted from controlled behavioral experiments into a reward function. The authors also argue the trained policy can run in reverse — serving as a computational tool for psychology by letting researchers manipulate social discomfort, spatial norms, and interaction strategies in a controlled way.
Real-world applications.
- Service and delivery robots navigating crowded sidewalks, hotel lobbies, or hospital corridors without forcing pedestrians to swerve.
- Assistive mobility devices and socially aware wheelchairs that respect personal space, especially when passing groups in conversation.
- Virtual agents and avatars in VR, telepresence, or games that behave in ways human observers find natural rather than mechanical.
- Crowd-management simulations and architectural layout testing, where the model can predict how a given pedestrian layout shapes comfort.
Industry relevance. Training cost is a practical selling point: the framework reaches socially aligned behavior within 10,000 steps in a 15 m × 15 m environment on a single NVIDIA 3090, which matters for teams without large-scale human trajectory datasets. The open-sourced Unreal Engine evaluation pipeline also gives product teams a standardized way to user-test navigation policies in VR before committing to hardware trials.
Future Directions
- Beyond simulation: This paper reports only simulated environments and a VR user study; no physical robot deployment is described, leaving the sim-to-real gap open.
- Generalizing the rules: The social field parameters (m_a = 0.321, n_a = 0.856, m_p = 0.438, n_p = 0.630, a = 0.285, b = 0.175, c = 1.430, K = 10.180) come from prior work; the paper does not report whether they transfer across cultures, age groups, or scenario types.
- Resolving a reported parameter discrepancy: The main text lists a discount factor of γ = 0.9 for the experiments while the appendix states γ is configured to 0.8; clarifying this would aid reproducibility.
- Richer social understanding: The static persons used here do not react to the agent — extension to dynamic, reciprocally reacting crowds, and to interaction states beyond the face-to-face and group formations tested, remains an open question.
- Wider validation: Participants were 30 university students and staff aged 18–29, so comfort ratings from older or more diverse populations are not reported.
Target Audience
Readers who will benefit most are researchers and practitioners in human–robot interaction and social navigation, reinforcement learning engineers interested in reward shaping from non-ML domain knowledge, and cognitive scientists who study personal space and social locomotion and want a computational testbed for their measurements. It is also a useful case study for teams building VR evaluation pipelines, since the dataset construction and rating procedure are described in enough detail to replicate.
Authors’ abstract
Navigating human-populated environments without causing discomfort is a critical capability for socially-aware agents. While rule-based approaches offer interpretability through predefined psychological principles, they often lack generalizability and flexibility. Conversely, data-driven methods can learn complex behaviors from large-scale datasets, but are typically inefficient, opaque, and difficult to align with human intuitions. To bridge this gap, we propose RLSLM, a hybrid Reinforcement Learning framework that integrates a rule-based Social Locomotion Model, grounded in empirical behavioral experiments, into the reward function of a reinforcement learning framework. The social locomotion model generates an orientation-sensitive social comfort field that quantifies human comfort across space, enabling socially aligned navigation policies with minimal training. RLSLM then jointly optimizes mechanical energy and social comfort, allowing agents to avoid intrusions into personal or group space. A human-agent interaction experiment using an immersive VR-based setup demonstrates that RLSLM outperforms state-of-the-art rule-based models in user experience. Ablation and sensitivity analyses further show the model's significantly improved interpretability over conventional data-driven methods. This work presents a scalable, human-centered methodology that effectively integrates cognitive science and machine learning for real-world social navigation.