Research
Realistic Synthetic Household Data Generation at Scale
Overview Research area: Embodied AI / robotics — synthetic data generation for household environments and long-term human–robot interaction. Technical level: Intermediate. The framework is easy to fol

- arXiv
- 2602.07243
- Published
- 2026-02-06
- Authors
- Siddharth Singh, Ifrah Idrees, Abraham Dauhajre
AI summary
Overview
Research area: Embodied AI / robotics — synthetic data generation for household environments and long-term human–robot interaction.
Technical level: Intermediate. The framework is easy to follow conceptually, but the evaluation leans on embedding-space statistics (SBERT, CLIP, mutual information, PCA, K-means, t-tests, ANOVA, Cohen's d) that assume some familiarity with machine learning evaluation practice.
Scope: The paper proposes and statistically validates a generative pipeline that produces 3D household environments and temporally extended human activity / human–robot interaction data, where personas shape environments and environments shape behavior.
What This Paper Is About
Robots designed for homes need training data that captures how people and their living spaces influence each other over time — a dishwasher in a kitchen implies dish-loading activity, and a persona who sketches implies art supplies and an art studio. Existing synthetic data generators treat environment generation and human activity generation as separate, decoupled processes, so neither one informs the other. This paper builds a generative framework that couples them, letting each side iteratively update the other so that the resulting household data is both semantically grounded and temporally consistent.
Key Contributions
- A framework for generating synthetic human activity and human–robot interaction data with temporal consistency across long time horizons, supporting co-authoring through natural language.
- A persona-driven indoor environment generation method that produces 3D household schematics reflecting household member characteristics, including personalized room assignment, work-from-home-based home offices, and persona-specific asset placement.
- Bidirectional coupling mechanisms that pass semantic information between the environment and activity generators iteratively, until convergence criteria are met.
- A statistical validation suite using multi-modal embeddings (SBERT for text, CLIP for visual floorplans), mutual information mediation analysis, intervention studies, and a sim-to-real comparison against the HOMER dataset and the Wang et al. synthetic dataset.
Main Findings
- Semantic alignment across components: SBERT cosine similarity was 0.68 ± 0.09 for Persona–Environment, 0.72 ± 0.07 for Environment–Behavior (the highest), and 0.61 ± 0.12 for Persona–Behavior. CLIP visual–textual similarity between house images and household descriptions was 0.74 ± 0.08.
- Environment acts as a mediator: The framework satisfies the mediation criterion MI(persona, env) + MI(env, beh) > MI(persona, beh), indicating the generated environment is an effective channel between persona traits and behavior patterns.
- Mutual information scales with data size: MI(persona,env)+MI(env,beh) was 0.27 ± 0.08 at 20 data points, 0.48 ± 0.07 at 50, 0.64 ± 0.06 at 100, and 0.70 ± 0.05 at 200; the corresponding PCA component counts were 14, 30, 45, and 51, and K-means cluster counts were 6, 15, 17, and 20.
- Iterative refinement improves coherence: Across five iterations, MI rose from 0.45 ± 0.09 to 0.85 ± 0.04 and cosine similarity rose from 0.58 ± 0.12 to 0.79 ± 0.06.
- Persona interventions produce measurable effects: With N=30 samples per condition, age teenager gave p<0.001, F=12.4, d=0.89; age retiree gave p<0.001, F=15.2, d=1.12; messy organization gave p=0.003, F=8.7, d=0.64; organized gave p=0.001, F=10.1, d=0.73; early sleep gave p=0.012, F=6.2, d=0.51; late sleep gave p=0.008, F=7.1, d=0.58. Effect sizes ranged from d=0.51 to 1.12, described as medium to large. t-SNE visualizations showed distinct, separable clusters per intervention type.
- Stronger alignment with real than with synthetic data: Cosine similarity against HOMER (self-reported activity schedules from 21 participants) was 0.60 (stated as 0.603 in the discussion text), which the authors classify as high similarity. Against the Wang et al. dataset (16 publicly available synthetically generated personas with 5-day activity sequences) it was 0.27, classified as low similarity.
- Computational cost: For a three-member household in a three-room environment over a single day, with bidirectional generation across five iterations, environment generation used 10 LLM calls and 50.00 s, human–robot interaction generation used 7 calls and 81.04 s (the paper attributes the longer time to a larger context window carrying cumulative interactions), and bidirectional generation used 5 calls and 19.00 s.
Methodology in Plain English
The system accepts natural language descriptions of household members (demographics, lifestyle, routines) plus high-level environmental constraints such as house type and room configuration. Rather than passing free-form text to every stage, the authors supply information in a structured way and pass intermediate outputs — including visual information about the floor plan — to later steps, maintaining a running "contextual memory" that tells the language model what task is being performed and what previous steps completed. The authors report this substantially reduces hallucination.
Four modules make up the pipeline: an Environment Schematic Generator that produces 3D layouts with semantic object selection and placement; a Human Activity and HRI Generator that synthesizes temporally consistent behavior sequences grounded in environmental affordances; a Bidirectional Influence Controller that mediates iterative information exchange; and a Universal Simulator Adapter that converts intermediate representations into formats usable by different simulators, making the framework simulator-agnostic.
Activity generation uses a three-stage hierarchical decomposition: first, structured activity sequences are generated from member profiles, environmental constraints, temporal parameters, and robot capabilities; second, natural dialogues between humans and robots are synthesized from those sequences, accounting for social and cultural factors; third, the simulator adapter converts the representation for platform-agnostic deployment. A least-to-most prompt tuning approach and a rolling window context mechanism keep events consistent across days and weeks.
On the environment side, the authors decouple the pipeline from any specific asset database — swapping databases works provided metadata covers asset description, dimensions, pivot point, and optionally an asset image. They correct common layout errors algorithmically, including nested rooms and disconnected rooms, and use LLM recommendations to decide which rooms connect (for example, open-concept dining and living areas with the wall removed).
The coupling loop works as follows: both generators start from the same persona information; the environment module then emits object inventories, spatial layouts, and affordance maps that constrain activity generation; the generated activities in turn inform object placement, room utilization, and environmental modification. The loop stops after a maximum iteration count or when user-specified convergence criteria are met — a weighted combination of environmental object density (objects per room), activity schedule granularity (total duration divided by number of activities), and semantic similarity (cosine similarity between environment and activity descriptions in SBERT embedding space). For evaluation, embeddings are standardized by removing names, using generic room names, and including specific asset descriptions, then reduced with PCA (n=10 components) or K-means (n=50 clusters) to create comparable feature spaces. Real-world alignment is assessed by encoding activity patterns as 18-dimensional hourly probability vectors covering 6 AM to 11 PM.
Why This Matters
The paper targets a concrete data bottleneck: companies building household robots need large, diverse datasets spanning family dynamics and home configurations, and collecting that data from real homes at scale is impractical. By coupling environment and behavior generation, the authors argue the resulting data better reflects the mutual influence that defines real households, which is what a robot must model to anticipate needs and adapt.
Real-world applications include:
- Cleaning robots that need to understand household routines and object placement to plan coverage.
- Assistive robots that anticipate human needs based on persona and time-of-day patterns.
- Smart home systems that adapt to changing family dynamics and configurations.
- Simulated testing of household smart devices at scale before physical deployment, enabled by the simulator adapter.
Industry relevance is explicit: the work originates from Amazon Lab 126 and the discussion states it was developed for the emerging indoor robotics industry, prioritizing deployable solutions and sim-to-real validation over academic novelty. The framework's natural language configuration also means non-experts can specify datasets by describing scenarios rather than writing code.
Future Directions
- Improving real-time adaptation capabilities of the generation pipeline.
- Expanding validation across more diverse household types and higher behavioral complexity — the authors note the current evaluation covers working families with school-age children and that performance varies across behavioral datasets.
- Reducing computational overhead from iterative refinement, which the authors list as a limitation alongside dependence on LLM performance and scaling challenges with complex household scenarios.
- Addressing generation failure modes the authors identify: interaction conflicts (simultaneous incompatible activities such as loud music during sleep) and hallucinated activities producing impossible scenarios.
- Developing standardized benchmark datasets for this problem domain.
Target Audience
Robotics and embodied AI researchers who need scalable synthetic training data; engineers at companies building household robots, smart home devices, or simulation testbeds; and researchers working on LLM-driven generative pipelines for environments or human behavior who want to see a concrete bidirectional-coupling architecture with an accompanying statistical validation recipe. Readers evaluating whether synthetic household data can stand in for real-world behavioral data will find the HOMER versus Wang et al. comparison the most directly relevant part.
Authors’ abstract
Advancements in foundation models have catalyzed research in Embodied AI to develop interactive agents capable of environmental reasoning and interaction. Developing such agents requires diverse, large-scale datasets. Prior frameworks generate synthetic data for long-term human-robot interactions but fail to model the bidirectional influence between human behavior and household environments. Our proposed generative framework creates household datasets at scale through loosely coupled generation of long-term human-robot interactions and environments. Human personas influence environment generation, while environment schematics and semantics shape human-robot interactions. The generated 3D data includes rich static context such as object and environment semantics, and temporal context capturing human and agent behaviors over extended periods. Our flexible tool allows users to define dataset characteristics via natural language prompts, enabling configuration of environment and human activity data through natural language specifications. The tool creates variations of user-defined configurations, enabling scalable data generation. We validate our framework through statistical evaluation using multi-modal embeddings and key metrics: cosine similarity, mutual information gain, intervention analysis, and iterative improvement validation. Statistical comparisons show good alignment with real-world datasets (HOMER) with cosine similarity (0.60), while synthetic datasets (Wang et al.) show moderate alignment (0.27). Intervention analysis across age, organization, and sleep pattern changes shows statistically significant effects (p < 0.001) with large effect sizes (Cohen's d = 0.51-1.12), confirming bidirectional coupling translates persona traits into measurable environmental and behavioral differences. These contributions enable development and testing of household smart devices at scale.