Skip to content
AI.info

Research

A Benchmark for Spatially Grounded Gesture Generation

Overview Research area: Computer vision and motion synthesis, specifically co-speech gesture generation for embodied agents (virtual avatars, social robots, AR assistants), with an emphasis on benchma

A Benchmark for Spatially Grounded Gesture Generation
arXiv
2610.03105
Published
2026-10-02
Authors
Anna Deichler, Rishabh Dabral, Fethiye Irmak Dogan, Anindita Ghosh, Jonas Beskow

AI summary

Overview

Research area: Computer vision and motion synthesis, specifically co-speech gesture generation for embodied agents (virtual avatars, social robots, AR assistants), with an emphasis on benchmark design and evaluation methodology.

Technical level: Advanced. The paper assumes familiarity with diffusion and flow-matching generative models, SMPL-X body representations, kinematic inverse kinematics, and distributional motion metrics such as Fréchet Gesture Distance.

Scope: The paper introduces a benchmark — dataset, task formulation, and evaluation protocol — for generating pointing gestures that are geometrically grounded in a 3D scene, and uses it to compare a flow-matching baseline, a retrieval-based system, and captured human motion.

What This Paper Is About

Existing co-speech gesture models are trained and scored on how natural or speech-synchronised their motion looks, using distributional metrics like Fréchet Gesture Distance and diversity. That is fine for beat and iconic gestures, but a pointing gesture aimed at the wrong object can score well on those metrics while failing completely as communication. The paper's goal is to build the missing infrastructure — data pairing gesture with 3D scene context, a task definition, and metrics — so that generative systems can be compared on whether their gestures actually indicate the intended object, not just on whether they look plausible.

Key Contributions

  1. A paired data release. Approximately 2K pointing-annotated clips drawn from MM-Conv's naturalistic VR dialogue (1,842 MM-Conv pointing clips with dominant-hand annotations) plus 1,138 clean single-target pointing clips from SGS-HSI, retargeted to shared actors and with no test clip sharing a source with training.

  2. A task formulation for referential gesture in continuous dialogue. Systems receive audio and a word-level transcript, a 3D target coordinate in pelvis-centred and world frames, the dominant pointing hand, a static SMPL-X body shape, and the room's scene graph. Ground-truth motion, apex frame, referent identity, and evaluator labels are withheld, so the system must decide when the gesture peak occurs without an oracle apex frame.

  3. An evaluation protocol with a ground-truth reference run. Temporal alignment (Score a) and spatial grounding (Score b) are measured separately and combined multiplicatively into a Spatio-Temporal Grounding score (STG), complemented by a crowdsourced naturalness study.

  4. A reproducible flow-matching baseline, MM-Conv-Flow, built from a 170.5 M-parameter, 12-layer Diffusion Transformer pretrained on BEAT2, adapted with target-conditioned residual branches, relational pointing losses, and an inference-time directional refinement, and evaluated on the held-out split.

  5. A validity analysis of the protocol, showing where the objective metrics disagree with human perception and with each other.

Main Findings

  • Spatial grounding favours the retrieval-based system. On the 100 held-out MM-Conv clips, RePointer scores Score b = 0.905 ± 0.195, above MM-Conv-Flow at 0.691 ± 0.291 and above captured motion at 0.767 ± 0.223. Its margin over MM-Conv-Flow is +0.214 (95% CI: +0.163, +0.272, bootstrap clustered over 59 source recordings), and over GT it is Δ = +0.140 (95% CI: +0.088, +0.196).

  • Temporal alignment separates the two systems. RePointer obtains Score a = 0.784 ± 0.278 versus 0.499 ± 0.402 for MM-Conv-Flow, with GT at 0.941 ± 0.210. In raw timing, median absolute apex errors are 0.40 s for RePointer and 0.92 s for MM-Conv-Flow. RePointer's advantage holds across all 806 tested combinations of the temporal tolerance τ and decay σ_t.

  • Combined STG is nearly tied between RePointer and captured motion. At the primary temporal parameters, RePointer reaches STG 0.719 ± 0.307 and GT 0.720 ± 0.265, while MM-Conv-Flow reaches 0.361 ± 0.365. Because the RePointer–GT ordering flips as the temporal parameters vary, the authors decline to draw a conclusion about RePointer relative to captured motion from STG.

  • Geometric precision does not translate into perceived naturalness. In the crowdsourced study, GT scored MOS 3.58 ± 0.10, MM-Conv-Flow 3.11 ± 0.10, and RePointer 2.95 ± 0.11. GT was rated significantly more natural than both generated conditions (W = 222, p_Holm < 10⁻⁶ for MM-Conv-Flow; W = 76, p_Holm < 10⁻⁹ for RePointer), but MM-Conv-Flow and RePointer did not differ significantly (W = 694, p_Holm = .071).

  • Motion-quality metrics also disagree with perception. RePointer obtained the lowest EMAGE FGD (2.364) versus MM-Conv-Flow (4.551) and held-out GT (4.738), and the highest L1 velocity diversity (0.760 versus 0.178 and 0.246) and beat alignment (0.380 versus 0.272 and 0.318), yet was not rated more natural.

  • Ablation shows the value of unfreezing and refinement. In development-stage evaluation on the full 455-clip held-out set with an earlier hold-based temporal scorer, Mode A (frozen adapter) reached Score b 0.315 and θ_eff 36.2°, Mode B (unfrozen) reached 0.593 and 19.5°, and Mode B plus the inference-time nudge reached 0.743 and 11.0°.

  • A qualitative failure mode of retrieval. In one observed case (clip 298), the retrieved anchor carried iconic object-interaction-like motion from the training corpus while the geometric adaptation redirected the arm toward the correct referent — good spatial grounding despite semantically inappropriate residual motion.

Methodology in Plain English

The authors start from MM-Conv, a corpus of 6.7 hours of dyadic VR dialogue across five AI2-THOR rooms, captured with SMPL-X motion at 30 fps, word-level transcripts, and per-room scene graphs containing 42–61 objects. Each of its 4,211 referring expressions was previously linked by hand to its referent object, and 49% of them are pronominal ("it", "that one"), meaning non-verbal cues matter. To build the benchmark they add a pointing annotation layer: a random-forest classifier using arm-kinematic, scene-context, and speech-timing features guided manual review, with human-assigned gesture tiers deciding inclusion. They hold out one room entirely; the fixed evaluation manifest is 100 items sampled from the held-out MM-Conv clips, stratified by speaker, spanning 75 unique targets across 21 object categories with targets 1.3–4.3 m away, azimuths from −176° to 177°, and elevations from −32° to 55°. Table 1 reports the split: 2,525 training clips (1,503 MM-Conv + 1,022 SGS-HSI) and 455 test clips (339 + 116).

For scoring, they avoid asking whether the motion "looks right" and instead ask two separate questions. Score a runs a frozen logistic-regression apex detector, trained on recordings disjoint from the test set, and compares the detected apex to the reference apex with a fixed tolerance of τ = 10 frames (0.33 s) and decay σ_t = 15 frames (0.5 s). Score b casts a ray from the wrist along the elbow-to-wrist direction at the detected apex and measures its distance to the target's bounding box, converting a miss into a "miss" angular-like error θ_eff = arctan(d_miss / d_centroid), mapped through a decay of σ = 20°. STG is simply the per-clip product of the two, averaged. Throughout, they are explicit that these are benchmark design parameters rather than empirically established perceptual thresholds.

The baseline, MM-Conv-Flow, is built in two phases. Phase 1 trains a 170.5 M-parameter, 12-layer Diffusion Transformer on BEAT2 via conditional flow matching, with motion represented as 300-dimensional SMPL-X features (25 joints × [3 positions, 6D rotations, 3 velocities]) at 30 fps over 180 frames (6 s), and speech conditioned through a frozen 306.8 M-parameter CSMP encoder using aligned audio and text data2vec features. Phase 2 injects the 3D target: a small MLP encodes unit direction plus log-distance with a learned hand embedding, and this is added either as a gated residual bias on layers 4–11 (Mode A, backbone frozen) or additionally through single-token cross-attention on layers 8–11 (Mode B). Training adds relational losses that align the forearm with the target direction and penalise insufficient arm extension, using a soft-min over time so only the best-pointing frames dominate. At inference, a ProjFlow-inspired directional nudge partially corrects the arm chain (wrist ×1.0, elbow ×0.5, shoulder ×0.25), gated by the model's own arm extension, with β = 0.5 and γ = 2 fixed on a training-room validation split.

The independent comparison system, RePointer, retrieves a full-body SMPL-X anchor from a memory of natural pointing clips indexed by target geometry, hand, timing, and apex ratio, then adapts it with an apex-preserving temporal warp, a target-ray inverse-kinematics step over the collar–shoulder–elbow–wrist chain, and a hold. Its IK objective explicitly optimises forearm–target alignment — the same quantity Score b reads.

Finally, a crowdsourced study renders 23 held-out clips under three conditions (GT, MM-Conv-Flow, RePointer) with identical audio, avatar, room, and camera, counterbalancing conditions with a cyclic Latin square. Sixty-two Prolific participants who passed both attention checks produced 1,410 ratings on a five-point naturalness scale, with 16 trials excluded for video load error.

Why This Matters

Impact on research. The paper reframes referential gesture generation as a task-success problem rather than a distribution-matching problem, and supplies the corpus, task, and metrics to make that reframing actionable. Its most pointed methodological message is negative: a system can beat captured human motion on a geometric grounding metric while being rated no more natural, and a system can have the best Fréchet Gesture Distance, diversity, and beat alignment without a perceptual advantage. That argues against collapsing gesture quality into a single ranking score.

Real-world applications:

  • Social robots and virtual avatars in shared physical space that must indicate objects to a human collaborator ("hand me that one").
  • AR assistants that overlay or animate an avatar pointing at a specific item in the user's environment.
  • Telepresence and mixed-reality avatars that need to reference objects in a shared room, not just gesticulate plausibly.
  • Assistive or instructional systems where a wrong referent is not merely less natural but genuinely misleading.

Industry relevance. Any group building speech-driven avatar animation now has a shared held-out split, released code-adjacent artefacts (SMPL-X motion, audio at 30 fps, transcripts, scene-graph JSON, and a sample-level annotation table at the Hugging Face address given in the paper), and a baseline whose scores future systems can be compared against directly.

Future Directions

  • Withholding the target coordinate. Both evaluated systems use the supplied target to improve grounding at inference time, so their scores partly reflect the strength of geometric operators rather than learned scene-aware generation. The authors identify the harder variant that requires inferring the referent from scene and utterance as the focus of future benchmark releases.

  • A referent identification study. The crowdsourced study measures naturalness only. Testing whether observers can actually identify the intended referent requires a dedicated item set in which each target has matched distractors at controlled angular separations, so chance level and difficulty are known; the current test set was sampled for speaker balance and target diversity instead.

  • Broader data. The benchmark rests on five rooms of a single VR corpus with one held-out room supplying all 100 test clips, and SGS-HSI's synthesised speech and constructed phrase–gesture alignment do not extend the naturalistic referential domain. The authors call for larger referential corpora spanning more scenes beyond indoor household rooms and more speakers.

  • Quantifying protocol sensitivities. The spatial score depends on estimated elbow and wrist positions, and sensitivity to pose-estimation or calibration errors remains unquantified; the temporal tolerance and decay are design choices rather than perceptual thresholds.

Target Audience

Researchers and engineers working on co-speech gesture generation, embodied conversational agents, and human–robot interaction who need a shared way to evaluate whether generated gestures refer to the right thing. It is also relevant to benchmark designers interested in how objective metrics can diverge from human perception, and to practitioners building avatar or robot animation pipelines who want a held-out evaluation set with a published baseline to compare against. A reader without background in motion synthesis or generative modelling will find the metric definitions accessible but the model sections demanding.

Authors’ abstract

Communication in shared space interweaves verbal and non-verbal signals, and pointing gestures anchor language to the environment: "put the cup on that one" is uninterpretable without the gesture that fixes the referent. Yet no common framework exists for evaluating whether generated gestures indicate their intended referent; distributional metrics reward a gesture aimed at the wrong object as long as it looks natural. We introduce a benchmark for spatially grounded gesture generation, comprising ~2K pointing-annotated clips from naturalistic VR dialogue with ground-truth 3D referents, a task in which systems must decide when, how and where to point within conversational speech, and a protocol that separates temporal alignment, spatial grounding and perceived naturalness. We also provide a flow-matching baseline, MM-Conv-Flow. Evaluating it alongside an independent retrieval-based system and captured human motion, we find that geometric grounding can exceed that of human pointing without any gain in perceived naturalness, showing that referential gesture quality must be measured along separate dimensions.

Read the original paper