Research
S$^3$IT: A Benchmark for Spatially Situated Social Intelligence Test
Overview Research area: Artificial intelligence — embodied social intelligence, spatial reasoning, LLM agent evaluation, and benchmarks. Technical level: Advanced. Scope: This paper introduces S$^3$IT

- arXiv
- 2512.19992
- Published
- 2025-12-23
- Authors
- Zhe Sun, Xueyuan Yang, Yujie Lu, Zhenliang Zhang
AI summary
Overview
Research area: Artificial intelligence — embodied social intelligence, spatial reasoning, LLM agent evaluation, and benchmarks.
Technical level: Advanced.
Scope: This paper introduces S$^3$IT, a procedurally generated benchmark in which an embodied agent must explore a 3D house, interview LLM-driven NPCs to learn their preferences and conflicts, and then produce a seating arrangement that satisfies interlocking physical and social constraints.
What This Paper Is About
Existing evaluations of social intelligence in AI either test models on static text narratives (disembodied social reasoning) or test physical task execution in 3D environments without meaningful social constraints. Neither assesses whether an agent can integrate the two — reasoning about people's preferences, relationships and conflicts while simultaneously respecting the physical layout of the space they occupy. The paper builds a benchmark around a seat-ordering task to measure exactly this combined ability, and finds that even the strongest current models fall far short of human performance, primarily because of deficits in spatial understanding.
Key Contributions
-
The S$^3$IT benchmark, a novel evaluation of an embodied agent's social reasoning and planning under intertwined physical and social constraints, built on a seat-ordering task that is described as NP-hard.
-
A procedurally extensible generation and evaluation framework that constructs test scenarios of controllable difficulty and diverse type, enabling fine-grained analysis of an agent's capabilities and limitations.
-
A three-phase testing pipeline (NPC preference extraction and summarization, environmental cognition, multi-constraint decision-making) with an automatic scoring system, including an oracle ablation that isolates spatial perception as the bottleneck.
-
A systematic evaluation of leading LLMs against a human baseline, identifying strengths in explicit rule-based conflict handling and limitations in spatial and embodied social reasoning.
Main Findings
-
Top model score: Gemini-2.5-pro was the best-performing model on the 70-question Test Set with an average score of 47.8, and was the only model to exceed 40 on the Embodied Preference dimension (40.6).
-
Full model ranking: Average scores were GPT-4o-mini 19.3, Claude-4.5 23.1, GPT-4.1-mini 26.8, GPT-4o 28.2, Doubao-1.5 28.3, GPT-4.1 29.3, o4-mini 41.4, GPT-5 42.7, o3 43.1, and Gemini-2.5-pro 47.8.
-
Large human gap: On the 10-question Human Test Set, humans averaged 84.7 versus Gemini-2.5-pro's 41.4. The gap was largest on the Embodied dimension (humans 90.6 versus 36.3).
-
Near-human conflict handling: Top models approached human performance on the Conflict dimension — o4-mini 89.5, o3 89.0, GPT-5 86.1, Gemini-2.5-pro 85.7, versus the human 89.8 — indicating proficiency with explicit, rule-based constraints.
-
Positive prioritization gap (PG) for all models: Every evaluated model produced a positive PG score that scaled with capability, meaning models could distinguish stronger from weaker preferences. The human PG was lower (1.7), attributed to a strategy of accommodating all preferences to maximize overall score.
-
Weakness on group-based social preferences: Models systematically underperformed on "Group of NPC" preferences relative to "Relation" and "Topic" categories. For example, Gemini-2.5-pro scored 78.7 on Relation, 57.8 on Group of NPC, and 63.7 on Topic; GPT-5 scored 76.1, 51.1, and 70.7 respectively. The paper attributes this to the multi-step, programmatic reasoning that group preferences require.
-
Spatial perception ablation: Supplying ground-truth perception produced dramatic gains — Doubao-1.5 rose from 28.3 to 75.5 average, o3 from 43.1 to 79.8, and Gemini-2.5-pro from 47.8 to 74.8, with embodied scores reaching 86.1, 80.1, and 78.5 respectively. This confirms the primary limitation is deriving structured 3D understanding from visual inputs, not reasoning itself.
-
Self-reflection helps: All three representative models tested showed consistent improvement across reflection iterations of the "generate-and-reflect" mechanism.
-
Human strategy observation: Human participants often used a heuristic "anchoring" approach, making local adjustments rather than systematic backtracking, and were sometimes trapped by flawed initial assumptions. Humans required an average of 165 minutes to complete the 10-problem test suite; the paper estimates the full 70-question set would take over 20 hours per person.
Methodology in Plain English
The researchers built a simulated small town of 59 residents organized into 11 families spanning up to 4 generations, each resident having a name, 3D appearance, demographics, job, workplace, residence, income, social relationships, interests, dominant hand, and three kinds of attributes: embodied preferences, social preferences, and social conflicts.
The task given to the agent under test (called the T-Agent) is to seat a group of NPCs in a 3D room. To do this, it must complete three phases:
-
Preference extraction: The T-Agent talks to each NPC in natural language, then distills the unstructured dialogue into a structured profile with a description and an intensity score (strong, medium, or low).
-
Environmental cognition: The T-Agent iteratively explores the 3D scene, choosing viewpoints to maximize information gain and fusing multi-view observations into structured environmental features that capture distances, orientations, and viewing angles between seats and elements such as windows, televisions, and air conditioners.
-
Decision-making: The T-Agent generates a seating plan plus a reflection report annotating which preferences are satisfied, then iteratively refines the plan using both static context (profiles, environmental features) and dynamic context (the previous plan and its reflection) until convergence or a fixed iteration limit.
Scenarios are generated by construction: the system starts with a random valid seating arrangement as ground truth, then derives constraints from it, guaranteeing every problem has at least one solution. The dataset contains 7,000 instances across 70 difficulty levels built from 5 scene templates, with each NPC typically having 1–5 preferences and 0–2 conflicts, each rated on a 3-point Likert-type scale. Rooms are randomized across 8 table shapes and textures and 6 chair shapes.
Scoring is automatic. A simulator checks embodied preferences by computing chair positions and distances to furniture and windows; a hand-crafted discriminator checks social preferences by examining immediate neighbors; conflicts count as avoided if immediate neighbors have no conflicts with the NPC. Scores are weighted, aggregated at the category level, and remapped through a polynomial curve to a continuous range of [0, 1] to penalize answers that satisfy only a small fraction of preferences within a category. Scores are then scaled to 0–100 for human review, with in-simulator visualizations showing NPCs in their seats, their preferences above their heads, relationship lines, and the T-Agent's perceived summary.
The Test Set consists of 70 questions, one sampled from each of the 70 difficulty levels, preserving the feature distribution of the full 7,000-question dataset. The Human Test Set is a 10-question representative subset chosen with an Integer Programming algorithm, evaluated by three recruited participants.
Why This Matters
The work reframes social intelligence evaluation: rather than asking whether a model can reason about people in the abstract, it asks whether a model can reason about people in a physical place. The finding that ground-truth perception lifts model scores from roughly 28–48 to roughly 75–80 suggests that the reasoning machinery in current LLMs is stronger than their ability to perceive and structure 3D space — a concrete pointer for where future effort should go.
Real-world applications:
-
Service robotics and assistive agents that must seat or arrange people in restaurants, clinics, or care facilities while respecting both physical layout and interpersonal dynamics.
-
Event and conference planning tools that allocate seating across large groups with family ties, colleague relationships, and known conflicts.
-
Workplace and office layout optimization where team composition, hierarchy, and proximity constraints interact.
-
Human-agent collaboration research testing whether agents can be trusted to make socially nuanced decisions in shared physical spaces.
Industry relevance: The benchmark targets the deployment gap for embodied agents entering human environments. Its procedural generation framework lets labs produce fresh test instances of controllable difficulty, which is directly useful to teams building embodied assistants, and the oracle ablation gives a clear diagnostic signal about whether a system's failures come from perception or from reasoning.
Future Directions
-
Continuous and more realistic interaction. The authors note that the discrete viewpoint set could be removed to require the T-Agent to plan continuous motion trajectories.
-
Uncooperative NPCs. NPCs could be configured to behave uncooperatively when queried about their preferences, adding a harder social dimension.
-
Broader human baselines. The human evaluation relied on only three experts, which the authors acknowledge limits generalizability to a broader population.
-
Scoring function calibration. The empirical polynomial curve used to penalize partial satisfaction may carry subjective bias, though the authors state it is applied consistently across categories so it does not change the rank order between LLMs and the human baseline.
-
Richer interaction protocols. Because S$^3$IT is interactive rather than narrative-based, varying the interaction process can influence performance, leaving room for more complex integration of T-Agents.
Target Audience
Researchers and engineers working on embodied agents, LLM agent evaluation, multimodal grounding, and human-agent collaboration. The paper is most useful to those designing benchmarks or building systems that must combine spatial perception with social reasoning. It is also relevant to practitioners in robotics and service automation who need to know where current model capabilities actually break down. A background in agent evaluation and 3D simulation helps, but the core argument and results are accessible to anyone familiar with LLM benchmarking.
Authors’ abstract
The integration of embodied agents into human environments demands embodied social intelligence: reasoning over both social norms and physical constraints. However, existing evaluations fail to address this integration, as they are limited to either disembodied social reasoning (e.g., in text) or socially-agnostic physical tasks. Both approaches fail to assess an agent's ability to integrate and trade off both physical and social constraints within a realistic, embodied context. To address this challenge, we introduce Spatially Situated Social Intelligence Test (S$^{3}$IT), a benchmark specifically designed to evaluate embodied social intelligence. It is centered on a novel and challenging seat-ordering task, requiring an agent to arrange seating in a 3D environment for a group of large language model-driven (LLM-driven) NPCs with diverse identities, preferences, and intricate interpersonal relationships. Our procedurally extensible framework generates a vast and diverse scenario space with controllable difficulty, compelling the agent to acquire preferences through active dialogue, perceive the environment via autonomous exploration, and perform multi-objective optimization within a complex constraint network. We evaluate state-of-the-art LLMs on S$^{3}$IT and found that they still struggle with this problem, showing an obvious gap compared with the human baseline. Results imply that LLMs have deficiencies in spatial intelligence, yet simultaneously demonstrate their ability to achieve near human-level competence in resolving conflicts that possess explicit textual cues.