Skip to content
AI.info

Research

Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks

Overview Research area: Human-Computer Interaction (HCI), specifically multimodal communication, convention formation, and augmented reality-mediated collaboration; also touches cognitive science and

Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
arXiv
2602.08914
Published
2026-02-09
Authors
Kiyosu Maeda, William P. McCarthy, Ching-Yi Tsai, Jeffrey Mu, Haoliang Wang, Robert D. Hawkins, Judith E. Fan, Parastoo Abtahi

AI summary

Overview

Research area: Human-Computer Interaction (HCI), specifically multimodal communication, convention formation, and augmented reality-mediated collaboration; also touches cognitive science and computational cognitive modeling (Rational Speech Act framework).

Technical level: Intermediate. The experimental designs and behavioral results are described in accessible terms, but the computational modeling component assumes some familiarity with probabilistic models of communication (RSA) and mixed-effects analysis.

Scope: The paper reports two studies of how people form shared linguistic and gestural conventions while instructing a partner in a physical block-assembly task, plus a computational model extending convention formation to multimodal settings.

What This Paper Is About

When people work together repeatedly, they invent shorthand—short words, shared references, and conventions—that make communication faster and more efficient. Most controlled studies of this phenomenon have looked only at text or 2D tasks, even though real physical collaboration involves speech plus hands.

This paper asks how those conventions form when the task is physical and communication is multimodal: how do speech and gesture each change over repeated assembly instructions, how does information get distributed across the two channels, and can a computational model capture those shifts?

Key Contributions

  1. A large online unimodal study (n = 98, paired into dyads) investigating natural language communication over repeated interactions in a tower-building task.
  2. An AR-mediated multimodal lab study (n = 40, 20 pairs) examining gesture and speech during repeated physical assembly, including a released multimodal dataset (audio transcripts and 4D gestures) and a custom desktop viewing application.
  3. A computational model and simulation that extends the Rational Speech Act framework to a multimodal lexicon, exhibiting behaviors aligned with the study findings.
  4. A flexible model design that the authors state is extendable to more synchronous and symmetric communication.

Main Findings

  • Online study accuracy improved: Reconstruction accuracy (F1 overlap between reconstructed tower and target silhouette) rose from a mean of 0.88, 95% CI [0.85, 0.90], in the first repetition to 0.98, 95% CI [0.96, 0.99], by the end (β = 0.92, t(54.84) = 6.22, p < .001).
  • Online instructions got shorter: Instructors used significantly fewer words (β = −8.53, t(36.9) = −9.58, p < .001) and sent fewer messages (β = −18.1, t(24) = −7.11, p < .001) across repetitions.
  • Vocabulary shifted from blocks to towers: Word frequencies that decreased most included "BLOCK," "TWO," "HORIZONTAL," "VERTICAL," and "ANOTHER"; those that increased most included "L," "C," "SHAPE," and "BLUE."
  • Dyad vocabularies diverged: Jensen–Shannon divergence between dyads' word-frequency distributions increased from R1 to R4 (0.080, 95% CI [0.041, 0.118], p = .004), indicating distinct linguistic mappings rather than a single shared vocabulary.
  • Abstraction level shifted: Annotators (ICC = 0.83, 95% CI [0.82, 0.84]) found a significant interaction between repetition and expression type (b = 0.53, t(47.5) = 4.8, p < .001). Mean block-level references dropped from 7.3 to 3.6, while tower-level references rose from 0.6 to 1.1.
  • Multimodal task success improved: Success rate rose from 74.03% in R1 to 98.36% in R4, with only one unsuccessful trial.
  • Multimodal efficiency improved: Repetition significantly reduced instruction length (β = −12.80, SE = 1.45, p < .001), time to first block (β = −2.99, SE = 1.10, p = .007), time to last block (β = −15.96, SE = 1.99, p < .001), and construction time (β = −20.19, SE = 2.73, p < .001).
  • Tower shape affected word count: Instructions for the L tower had significantly fewer words (β = −22.70, SE = 7.70, p = .003), while instructions for the TREE tower had significantly more (β = 26.05, SE = 7.70, p = .001), though the gap narrowed across repetitions.
  • Tower words appeared early: 86.7% of Instructors used at least one tower word in R1. Tower words were used 82 times as "C" and 29 times as "L"; TREE was called "Tree" 4 times, "T" 36 times, and "Cross" 21 times.
  • Tower instructions took more time later: The proportion of time spent on tower instructions increased from 11.93% in R1 to 27.29% in R4 (permutation test with 10,000 data shuffles, p = .015).
  • Tower words moved earlier in messages: The relative position of tower words decreased across repetitions (β = −0.44, SE = 0.096, p < .001), with a significant orthogonalized quadratic term (β = 0.071, SE = 0.019, p < .001) indicating the rate of change slowed.
  • Builders changed strategy too: In R1, 19 out of 20 Builders placed blocks one by one on the grid; in R4, 11 Builders constructed the entire tower before placing it.
  • Gestural abstraction did not significantly shift in proportion: There were 78 tower gestures (41 static, 37 stroke-based), 36 of them bimanual, but the analysis found no significant change in the proportions of block versus tower gestures across repetitions. The authors interpret this as participants using tower gestures in R1 to establish a convention, then increasingly referencing tower shapes with words alone.
  • The model worked: Simulation results indicated the model successfully acquired abstract tower programs (RQ1) and captured modality-dependent behavioral shifts across repetitions (RQ2).

Methodology in Plain English

The researchers ran two studies to watch how communication changes when the same ideas have to be explained again and again.

Study 1 (online, unimodal). 146 participants were recruited and paired into 73 dyads; 24 dyads were excluded for failing preregistered criteria (at least 75% reconstruction accuracy on at least 75% of trials, or self-reported confusion or non-fluency in English), leaving 98 participants. One person was the Instructor and saw a target scene of two towers; the other was the Builder and saw an empty grid. Instructors typed step-by-step instructions (one message up to 100 characters per turn) and Builders placed blocks. Each scene had two towers, each tower built from four domino-shaped blocks (two vertical, two horizontal). Three unique towers appeared in randomized order across four repetitions, giving twelve trials. Sessions lasted 30–50 minutes, with a minimum compensation of $5.00 plus up to $3.00 performance bonus.

Study 2 (AR lab, multimodal). 40 participants (20 pairs: 23 women, 14 men, 3 non-binary; ages 18–44, M = 23.68, SD = 5.05) were assigned to Instructor or Builder roles in separate rooms, each wearing a Meta Quest 3 headset. The Instructor viewed a 3D virtual twin of the target tower and recorded spoken plus gestural messages of up to 2 minutes. The system transmitted audio and 24-joint hand keypoints to the Builder and images of the Builder's assemblies to the Instructor. This isolated speech and gesture and removed other cues such as eye gaze and facial expressions. Builders replayed messages and assembled three-block LEGO towers (C, L, and TREE) on a 2×2 grid across twelve trials. Sessions lasted about 1 hour with $25 compensation.

Analysis. Speech was transcribed with Whisper (large-v2), with 11 of 256 sentences manually corrected. Two authors annotated references in R1 and R4 for abstraction level (block or tower), clarity (clear or ambiguous), and information type (shape, position, orientation); the same authors annotated gestures, using a collaborative coding approach rather than computing inter-rater reliability.

Modeling. The authors extended the Rational Speech Act framework to a multimodal lexicon mapping symbols to utterances and gestures, allowing for redundant, complementary, and language-only signal combinations, and assuming the lexicon returns continuous real values.

Why This Matters

Impact on research. Prior work on convention formation focused on unimodal text or 2D graphical tasks. This paper pushes the paradigm into physical, multimodal space, shows that linguistic and gestural abstraction co-evolve, and offers an extended RSA-style model as a starting point for studying multimodal conventions computationally. The released dataset of audio transcripts and 4D gestures supports further work.

Real-world applications:

  • AI-powered wearables and smart glasses that infer intent from egocentric audio and hand tracking, matching the study's egocentric presentation choice.
  • Robots and embodied agents that need to form shorthand conventions with a human partner across repeated physical tasks rather than only interpreting one-off commands.
  • AR instruction systems for assembly, repair, or construction where guidance can be progressively compressed as the user learns.
  • Systems that communicate intent through embodied, spatially aligned gesture output as well as speech.

Industry relevance. The work targets the gap between one-shot multimodal intent recognition and the repeated interactions that characterize real collaborative work in manufacturing, logistics, and training. Companies building AR headsets, collaborative robots, or gesture-aware assistants have a stake in whether agents can track and reuse conventions instead of re-explaining everything each round.

Future Directions

  • More synchronous and symmetric communication: The authors state their flexible model design is extendable beyond the Instructor-to-Builder setting studied here.
  • Richer visual feedback: The authors note that pilot testing confirmed 2D images were sufficient for this task, but that more complex towers or physical tasks may require live, photorealistic 3D renderings.
  • Explaining the gesture finding: Gestural abstraction proportions did not change significantly across repetitions (unlike speech), leaving open the question of how gestural conventions function differently from linguistic ones over time.
  • Convention-aware agents in the physical world: Applying the model's modality-preference shifts in agents that interact with humans during iterated physical collaboration, which the authors frame as the motivation for the modeling work.

Target Audience

HCI researchers studying multimodal interaction, augmented reality collaboration, and computer-supported cooperative work; cognitive scientists studying convention formation, lexical entrainment, and gesture; and researchers building computational models of communication (particularly RSA variants) who want to extend them beyond text. Designers of AR/VR collaboration tools, wearables, and human-robot interaction systems will also find the behavioral findings directly applicable.

Authors’ abstract

A quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner's hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world.

Read the original paper