Skip to content
AI.info

Research

SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration

Overview Research area: Human-Computer Interaction, with ties to computer vision, natural language processing, and mixed-reality systems. Technical level: Intermediate. The paper assumes some familiar

SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration
arXiv
2511.02560
Published
2025-11-04
Authors
Dan Bohus, Sean Andrist, Ann Paradiso, Nick Saw, Tim Schoonbeek, Maia Stiber

AI summary

Overview

Research area: Human-Computer Interaction, with ties to computer vision, natural language processing, and mixed-reality systems.

Technical level: Intermediate. The paper assumes some familiarity with egocentric computer vision datasets, speech recognition pipelines, and mixed-reality headsets, but the dataset itself is described in accessible terms.

Scope: The paper introduces SigmaCollab, a multimodal, egocentric dataset of 85 interactive sessions (13 hours, 45 minutes, and 11 seconds) in which untrained participants were guided by a mixed-reality AI agent through physical procedural tasks.

What This Paper Is About

Most existing egocentric and procedural-task datasets record a single person performing an activity alone, or record two humans interacting with each other. Neither setup captures what happens when a person collaborates with an AI system in the physical world. The authors address this by collecting data from people performing real hands-on tasks while being guided step-by-step by an actual working mixed-reality assistant application called Sigma, producing a resource aimed specifically at interaction and coordination challenges rather than only perception challenges.

Key Contributions

  1. The SigmaCollab dataset: 85 task execution sessions from 21 participants, spanning 13 hours, 45 minutes, and 11 seconds, containing 1,583 attempted sub-step executions and 3,296 user utterances, with multimodal streams including color, depth, and stereo grayscale camera views, head pose, eye gaze, hand pose, and audio.

  2. An application-driven, interactive data collection method: Rather than scraping video, simulating environments, or staging human-human interactions, the authors collected data by having participants use a real open-source AI application to pursue meaningful physical goals, which the authors argue yields more ecologically valid data.

  3. Rich post-hoc annotations: Manual transcriptions of user speech, word-level timestamps computed via force alignment for both user and system utterances, four-way task success labels for every session, and gaze post-processing that marks when participants were reading the virtual instruction panel and projects gaze points into the color, grayscale, and depth image spaces.

  4. Eight expert demonstration sessions: one for each of the eight tasks, performed by an author highly familiar with the tasks, released alongside the participant data for one-shot learning and prior-information uses.

Main Findings

  • Dataset scale and composition: The 21 participants attempted 95 task execution sessions; 10 were excluded due to application crashes and performance issues, leaving 85. The dataset spans 13 hours, 45 minutes, and 11 seconds, contains 1,583 attempted sub-step executions (1,570 after removing abandoned or interrupted sub-steps), and 3,296 user utterances.

  • Task success rate: Overall, 60 sessions were correctly completed, 8 incorrectly completed, 12 abandoned, and 5 were system failures, producing a task success rate of 75.0% across the dataset. Per-task success rates ranged from 40.0% (Whiskey) and 50.0% (Mojito) to 100.0% (Hard-drive and Margarita).

  • Runtime speech recognition is noisy: The Silero voice activity detector had an average session-level detection error rate of 12.4%, and the Whisper tiny.en recognizer had an average session-level word-error-rate of 20.2%. The authors state these numbers confirm significant runtime errors and motivate the manual transcripts.

  • Wide variability in interaction length: The shortest attempted task execution was 2 minutes and 1 second and the longest 26 minutes and 11 seconds. Sub-step executions ranged from 2.16 seconds to 7 minutes and 53 seconds, averaging 28.6 seconds.

  • A new interaction phenomenon surfaced by the method: The authors report that participants often talk to themselves throughout task execution, which raises the challenge of detecting self-talk, that is, deciding which utterances should be responded to and which should not.

  • Stream configuration and achieved rates: The color camera is 896 × 504 pixels at 24 bpp with a target of 15 Hz and an actual average of 14.91 Hz; the long-throw depth camera is 320 × 288 pixels at 16 bpp with a target of 5 Hz and actual 4.98 Hz; both 640 × 480 8 bpp front grayscale cameras ran at 13.64 Hz against a 15 Hz target; head pose plus eye gaze ran at 28.37 Hz against a 30 Hz target; hand pose (26 joints per hand) ran at 20.01 Hz against a 20 Hz target; audio is 1-channel, 32-bit floating-point PCM at 16.00 kHz against a 16 kHz target.

  • No fixed splits: The authors state that they do not define specific train/val/test splits, since the dataset is intended as a testbed that may be out-of-distribution for trained models, but they note that splits can be stratified by dataset section, participant id, task type, step type, task success, or combinations of these.

  • Expert demonstrations are shorter: The eight expert demonstration sessions are all successful and generally shorter than the corresponding participant sessions; for example, the expert skateboard demonstration took 10:26 and the expert notebook demonstration took 10:07.

Methodology in Plain English

The authors took an existing open-source mixed-reality task assistant, Sigma, which runs on a HoloLens 2 headset and talks users through procedural tasks, and used it as the instrument for data collection. They set up workstations for eight physical tasks in a lab: making coffee with a Nespresso Pixie machine, replacing a hard-drive in a PC, installing wheels on a skateboard, making a pin-back button, making a notebook with a binding machine, and preparing three mocktails (Margarita, Mojito, Whiskey Sour).

Twenty-one participants were recruited from within the organization, given a five-minute orientation slide deck, fitted with the headset, and eye-calibrated. Each then performed up to six randomly chosen tasks, in a stratified random order across six workstations, with the option to abandon any task. Sessions lasted up to ninety minutes and participants were compensated 75 USD. A researcher observed remotely through live data visualizations and only intervened when the application crashed or a participant was fully blocked after an unrecoverable mistake.

During the study the system streamed sensor data from the headset to a desktop server, which ran speech recognition, interaction management, and speech synthesis, and persisted the streams. The system classified each user utterance into one of four classes (question, no-question, next-step, or step-navigation) using an LLM with a prompt that included the task recipe, current step and sub-step, dialog history, and an egocentric grayscale image captured at the start of the utterance. Three Azure GPT deployments were used over the course of collection: a GPT-4o deployment in a global data zone region, then a GPT-4o-mini deployment in a US data zone region, then GPT-4o again in a US data zone region for the last third of the dataset. After collection, the team manually transcribed speech, computed word-level timings by force alignment, labeled each session's outcome, and derived gaze annotations.

Why This Matters

The work targets a gap the authors identify between static perception benchmarks and the messy realities of real-time collaboration, including timing, proactive intervention, grounding of references, and detecting user cognitive states such as frustration and confusion. Because the collection uses an open-source application, models developed or evaluated on this data can be deployed back into the same application and tested in live interactions, which the authors argue lets researchers measure end-to-end effects on task performance and user satisfaction rather than only offline accuracy.

Real-world applications this could inform:

  • Remote and in-person task assistance, where an AI guides a worker through assembly, repair, or setup while the worker's hands are occupied.
  • Mixed-reality industrial or field procedures, such as hardware replacement or equipment maintenance, where correct end states matter and mistakes can be unrecoverable.
  • Training and onboarding systems that need to detect when a trainee is confused, off-track, or talking to themselves rather than to the system.
  • Assistive robotics and embodied agents that must coordinate actions with a human in a shared physical space.

Industry relevance: The dataset was produced at Microsoft Research using Microsoft HoloLens 2 hardware, the open-source Sigma application, and Azure GPT-4o and GPT-4o-mini deployments, so it is directly tied to commercial mixed-reality and cloud model infrastructure. The paper is released under CC BY 4.0 and the dataset is available at https://github.com/microsoft/SigmaCollab.

Future Directions

  • Building benchmarks: The authors state they plan to use the dataset to construct a set of benchmarks for physically situated collaboration that complement existing egocentric computer vision benchmarks, focused on interaction-related challenges such as timing, proactive interventions, grounding, and detecting cognitive states like frustration and confusion.

  • Self-talk detection: The observation that participants frequently talk to themselves raises the open question of determining which utterances an agent should respond to, a competency the authors say has not been addressed in the community.

  • Closing the loop between offline evaluation and live deployment: Because the Sigma application is open source, an open question is how models trained or evaluated on this data change user behavior and task-level outcomes once deployed, and how those effects should be measured.

  • Extending the data: The authors describe successive cycles of refinement in which deploying models back into the application surfaces new questions and challenges, and they invite the community to collect additional data with the same tooling.

Target Audience

Researchers and practitioners working on egocentric computer vision, multimodal interaction, mixed-reality assistance, and human-AI collaboration, particularly those who need data collected under realistic interactive conditions rather than staged or scripted ones. It is also relevant to teams building embodied or task-assistive AI systems that want a benchmark for interaction-level competencies such as grounding, timing, and modeling user intent. Readers looking for very large-scale datasets should note the authors' own framing that the dataset is relatively small at approximately 14 hours, and that its value lies in its application-driven and interactive nature rather than its size.

Authors’ abstract

We introduce SigmaCollab, a dataset enabling research on physically situated human-AI collaboration. The dataset consists of a set of 85 sessions in which untrained participants were guided by a mixed-reality assistive AI agent in performing procedural tasks in the physical world. SigmaCollab includes a set of rich, multimodal data streams, such as the participant and system audio, egocentric camera views from the head-mounted device, depth maps, head, hand and gaze tracking information, as well as additional annotations performed post-hoc. While the dataset is relatively small in size (~ 14 hours), its application-driven and interactive nature brings to the fore novel research challenges for human-AI collaboration, and provides more realistic testing grounds for various AI models operating in this space. In future work, we plan to use the dataset to construct a set of benchmarks for physically situated collaboration in mixed-reality task assistive scenarios. SigmaCollab is available at https://github.com/microsoft/SigmaCollab.

Read the original paper