Research
ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset
ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset Overview Research area: Social robot navigation, human-robot interaction, and pedestrian trajectory prediction — specifically, large-
- arXiv
- 2607.21964
- Published
- 2026-07-24
- Authors
- Shashank Rao Marpally, Allan Wang, Atharva Ghotavadekar, Renato Alexandre Ribeiro, Nhat Le, Pilar Bachiller-Burgos, Pranav Goyal, Subham Agrawal, Yasuhiro Nitta, Howard Ziyu Han, Daeun Song, Masaki Kuribayashi, Kohei Uehara, Xiyue Wang, Yangzhe Kong, Duc M. Nguyen, Amirreza Payandeh, Gerardo Pérez-González, Alejandro Torrejón-Harto, Jeeho Ahn, Tisha Jain, Andrew Stratton, Elvin Yang, Jorge de Heuvel, Nico Ostermann-Myrau, Sai Anudeep Sajja, Mithilya Raj, Daisuke Sato, Gaston Rouquette, Nikolas Martelaro, Maki Sugimoto, Hironobu Takagi, Chieko Asakawa, Maren Bennewitz, Aaron Steinfeld, Xuesu Xiao, Christoforos Mavrogiannis, Harold Soh
AI summary
ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation DatasetOverview
Research area: Social robot navigation, human-robot interaction, and pedestrian trajectory prediction — specifically, large-scale multi-modal dataset collection and benchmarking for these fields (arXiv:2607.21964v1 [cs.RO], posted 24 Jul 2026; CC BY 4.0; a preprint of a manuscript under review at The International Journal of Robotics Research).
Technical level: Intermediate. The paper is a dataset-and-benchmark contribution; it is readable without deep math, but assumes familiarity with concepts such as social navigation, teleoperation, odometry, LiDAR point clouds, homography, and trajectory prediction.
One-sentence scope: The authors assemble and release ACME, a cross-cultural, multi-embodiment social-navigation dataset collected by 8 teams in 5 countries on 7 robot embodiments, containing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data.
What This Paper Is About
Robots that move among people need to follow unwritten social rules — how close is too close, which side of a corridor to use, how to yield — and those rules differ by culture, location, robot shape, and crowd density. Existing datasets for learning these behaviors are typically collected in one country, with one robot, and rarely include explicit robot-to-crowd communication, so models trained on them generalize poorly. ACME's goal is to provide a large, diverse, multi-modal real-world dataset of goal-directed social navigation and human movement, together with benchmarks, so that navigation policies and pedestrian trajectory predictors can be trained and tested across cultural and embodiment variation.
Key Contributions
- A large, geographically and morphologically diverse dataset. The authors state it is the largest (by duration) and most diverse (by geographical location and robot embodiment) human-demonstrated social-navigation and pedestrian trajectory prediction dataset, collected at 8 locations across 5 countries on 7 robot embodiments.
- The largest human-labeled Bird's-Eye View (BEV) pedestrian trajectory prediction dataset, including camera calibration information so that pixel trajectories can be transformed into metric ground-plane coordinates.
- Qualitative and quantitative comparison against prior datasets in terms of scenario complexity and pedestrian trajectory characteristics, plus benchmark results comparing state-of-the-art vision navigation and pedestrian trajectory prediction models on ACME.
- Usability packaging: human-readable synchronized multi-sensor data, raw ROS2 bag files, BEV video, human-verified pedestrian trajectories, scenario tags, semantic scene tags, and anonymization-support detections (the paper states all software developed for the dataset will be open-sourced).
Main Findings
- Scale: ACME contains 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data, collected across 8 sites in 5 countries on 3 continents by 8 data-collection teams; 7 of the 8 locations are university campuses, with Miraikan being a science museum.
- Robot diversity: The paper describes 7 robot embodiments and characterizes them as 5 visually distinct robots — 2 quadrupeds (NUS: Unitree GO2 EDU; UBonn: Unitree GO1), suitcase-style differential-drive robots (Miraikan, CMU, Keio: Cabot), 1 rover-style differential-drive robot (GMU: AgileX Scout Mini), and 2 mobile service robots (UMich: Stretch 2; UEx: Shadow robot).
- Per-site onboard duration (hours): NUS 6.54, Miraikan 7.53, CMU 1.56, Keio 1.35, UMich 3.2, UEx 3.29, GMU 3.45, UBonn 2.43.
- Per-site BEV duration (hours): NUS 10.01, Miraikan 12.39, CMU 2.65, Keio 3.84, UMich 5.10, GMU 2.30, UBonn 7.27; the UEx subset has no BEV data.
- Per-site BEV trajectory counts: NUS 31,848, Miraikan 21,322, CMU 1,987, Keio 5,095, UMich 3,728, GMU 3,186, UBonn 4,892 (UEx not applicable).
- Cultural coverage: ACME spans 5 countries, compared with 2 for Bi³ and 1 for each of the other trajectory datasets listed (ETH, UCY, Edinburgh, VIRAT, Town Centre, Grand Central, CFF, SDD, L-CAS, WildTrack, JRDB, ATC, THÖR, Crowd-Bot, TBD, SiT, THÖR-MAGNI).
- Cultural motivation: the paper cites a large-scale study (Sorokowska et al.) reporting significant variance in preferred social, personal, and intimate distances across country, age, and gender, and notes that preferred walking side varies — the United States, Spain, and Germany favor the right side, while Singapore, Japan, and India favor the left.
- Scenario distribution (tag counts): Against Traffic 2,234; Passing Conversational Groups 1,854; Open Area 1,781; Wide Corridors 1,468; With Traffic 1,196; Narrow Corridors 864; Intersections 713; Blind corner 536; Navigating through large crowds 352; Line Formation (like queues) 328; Entry/Exit 306; Overtaking Pedestrian(s) 251.
- Goal-oriented data collection: Compared with SCAND and MuSoHu, ACME trajectories are shorter and more localized — the paper shows distributions of individual trajectory duration and distance traversed as evidence that ACME captures goal-directed scenarios rather than long unconstrained wandering.
- More challenging scenes: Density and traversability analysis from the egocentric view (comparing ACME with SCAND and MuSoHu) is presented as evidence that ACME captures more challenging scenarios and a broader distribution of pedestrian behavior than previous datasets.
- Robot speech: A subset of ACME includes robot speech data — specifically the NUS, UEx, Miraikan, and Keio subsets — using a restricted phrase set: "Excuse Me", "Attention, Robot Here", and "Please Give Way".
- Multi-modal sensing: All trajectories include locally accurate robot poses (odometry), egocentric RGB, 3D LiDAR point clouds, and sensor calibration. Sensing hardware varies by site (e.g., RealSense D435i, Hesai XT-16, Velodyne VLP-16, Zed 2 / Zed 2i, Ricoh Theta z1, Robosense BPearl & Helios, GoPro, Logitech C920). Miraikan, CMU, and Keio additionally collected accurate localization with pre-built maps; GMU and UEx collected egocentric 360 RGB.
- Privacy handling differs by institution: no anonymization (GMU, CMU); face blurring (NUS, UMich, Keio); full-body segmentation (UEx, Miraikan, UBonn). To compensate for lost information, bounding box, segmentation mask, and keypoint detections (Yolov8 with ByteTrack) are released.
- Benchmark results are not reported in the available text. The abstract and contributions state that state-of-the-art vision navigation and pedestrian trajectory prediction models were benchmarked on ACME, but the numerical benchmark outcomes are not present in the provided excerpt.
Methodology in Plain English
Instead of collecting everything in one place with one robot, the authors distributed the effort: eight research teams in five countries each followed a shared set of recording rules. Each team picked busy locations — corridors, intersections, blind corners, doorways, open areas — and times of day when crowds were likely, so that every robot trajectory would encounter at least one person.
Robots were driven by human teleoperators (a "Wizard of Oz" setup, where the operator stays out of sight so pedestrians believe the robot is autonomous; this does not apply to the suitcase robots, which are physically guided). Teleoperators fixed a navigation goal before starting, and marked "Begin", "Discard", and "End" signals so each recorded trajectory corresponds to a meaningful, goal-directed path rather than aimless wandering. Onboard sensor data was recorded as ROS2 bags, then anonymized and downsampled into human-readable form at 4 Hz.
For human movement, overhead bird's-eye-view cameras recorded pedestrians. Initial trajectories were produced automatically with ByteTrack, then human annotators verified them at 10 Hz using a custom tool offering Relabel, Missing, Break, Join, Delete, Disentangle, and Undo corrections. Because pixel coordinates are not physically meaningful, the authors computed homographies to map image coordinates to ground-plane metric coordinates — using poster-sized AprilTags for Miraikan, Keio, GMU, and UBonn (and, for NUS, upright tags seen simultaneously by robot and BEV cameras), and measured stationary ground-truth keypoints such as tile intersections for CMU and UMich. Metric trajectories were cleaned (dropping segments shorter than 3.2 s, fixing jumps with instantaneous velocity above 5 m/s by interpolation, smoothing with a window size of 5, labeling stationary segments where average speed over 8 time steps is below 0.5 m/s, and limiting trajectories to within ±25 m, ±25 m), then downsampled to 2.5 Hz and organized in ETH/UCY-compatible formats. Finally, annotators tagged robot trajectories with scenario labels and segmented BEV scenes into semantic regions using the same location categories.
Why This Matters
Impact on research. Social navigation and trajectory prediction models are usually trained and evaluated on narrow distributions — often the comparatively small, predominantly outdoor ETH/UCY datasets — which limits generalization. ACME offers a single resource spanning cultures, environments, embodiments, and explicit robot speech, enabling studies of how context changes appropriate behavior, and it supplies human-verified metric trajectories rather than relying purely on error-prone automated tracking.
Real-world applications:
- Service and delivery robots in mixed public spaces — malls, airports, campuses, and sidewalks where robots must negotiate crowds, queues, doorways, and narrow corridors.
- Assistive navigation — the Miraikan and other partners work on guiding users through crowded public environments such as museums, where yielding and communication behavior matters for safety and trust.
- Museum and campus guide robots — the suitcase-style robots used at Miraikan, CMU, and Keio represent platforms meant to accompany people through populated spaces.
- Robot dogs and rovers for security, inspection, and logistics — quadrupeds (Unitree GO2, Unitree GO1) and rovers (Scout Mini) have different navigation affordances and social presence, and the dataset captures how people respond to them differently.
Industry relevance. Companies deploying robots across countries cannot assume that a policy tuned in one locale transfers to another; hallway-side preferences, comfortable passing distances, and yielding conventions differ. ACME's scenario tags let engineers filter data for specific situations (high-density crowds, intersections, conversational groups) to train or evaluate policies for targeted deployments, and its ROS2 bag files plus preprocessed 4 Hz data make it practical to integrate into existing robot software stacks.
Future Directions
- Cross-cultural generalization studies: the dataset makes it possible to train on some locations and test on others, directly probing how well navigation policies and trajectory predictors transfer across cultural and geographical contexts — a question the paper raises but does not fully resolve here.
- Learning communication policies: the robot-speech subset (NUS, UEx, Miraikan, Keio) with utterance timing and content opens the question of when and how a robot should speak to proactively influence crowd behavior, which remains under-explored.
- Moving beyond ETH/UCY for trajectory prediction: with large, human-verified, metric-coordinate BEV trajectories, future work can benchmark and improve predictors on broader and denser human behavior distributions.
- Extending coverage and standardizing evaluation: the paper's motivation suggests open questions about adding more countries, more embodiments, and reconciling the conflicting social norms that arise when multiple context-dependent rules apply at once in the same scene.
Target Audience
Robotics and human-robot interaction researchers working on social navigation, crowd navigation, and pedestrian trajectory prediction; dataset and benchmark builders interested in multi-site, multi-modal collection methodology; and industry engineers building service, delivery, guide, quadruped, or rover robots intended to operate around people in public spaces. Readers looking for model architectures or algorithmic novelty will find this is primarily a data, annotation-protocol, and benchmarking resource.
Authors’ abstract
Understanding how robots and humans move in shared spaces is essential for designing effective social robot navigation policies and predicting human behavior. However, existing datasets often lack the diversity needed to capture differences in culture, geography, and human-robot interaction-factors that strongly shape appropriate social behavior. To address this gap, we introduce ACME: A Cross-cultural, Multi-Embodiment dataset for social navigation. A large-scale data collection effort across 8 sites in 5 countries, using 7 robot embodiments, ACME is a large and diverse multi-modal dataset aimed at advancing social navigation research, providing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data. Unlike prior datasets, it focuses on capturing goal-driven social navigation behavior in complex social scenarios with explicit robot-crowd interaction through robot speech. To facilitate learning navigation policies and predicting pedestrian trajectories, ACME provides 3D and 2D scene features, odometry, interaction information, and human-annotated pedestrian trajectory labels. We make ACME easy to use by providing both human-readable data for each sensor modality as well as raw binary data. Our qualitative and quantitative analyses show that our dataset captures more challenging scenarios and a broader distribution of pedestrian behavior than previous datasets.