Research
REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening
Overview Research area: Embodied conversational AI and robotics — audio-driven generation of reactive listener facial motion, with deployment on a physical humanoid robot. Technical level: Advanced. T

- arXiv
- 2609.33095
- Published
- 2026-09-27
- Authors
- Peizhen Li, Longbing Cao, Yang Zhang
AI summary
Overview
- Research area: Embodied conversational AI and robotics — audio-driven generation of reactive listener facial motion, with deployment on a physical humanoid robot.
- Technical level: Advanced. The paper assumes familiarity with transformer attention, learned gating, latent-variable residual models, motion-retargeting pipelines, and standard facial-motion benchmark metrics.
- Scope: The paper introduces REALM, a two-stage generative framework that predicts a coarse listener head-motion trajectory from speaker audio and listener history, then adds audio-conditioned stochastic residuals in the expression subspace, and demonstrates the result on the Ameca robot with a 25-participant perceptual study.
What This Paper Is About
Generating the facial motion of a listener — not the speaker — is hard because a person's response to a conversational cue does not happen at the same instant as the cue, and because some facial events (brief expressions, blinks) are locally unpredictable while the overall head trajectory is relatively smooth. REALM addresses this by combining a delay-centered attention prior over speaker audio with adaptive gating against the listener's own recent motion, then splitting generation into a coarse trajectory and a stochastic expression-only refinement. The goal is listener motion that stays continuous with the listener's own past behavior while still reacting appropriately to the speaker, and that can be retargeted onto physical robot hardware.
Key Contributions
- Delay-aware reactive fusion. A Reactive Gated Speaker–Listener Fusion module that pairs a shifted ALiBi-style attention bias, centered on a nominal reaction delay τ, with a learned gate g_t ∈ [0, 1] that adaptively weights aligned speaker context against encoded listener history.
- Expression-specific stochastic refinement. A coarse-to-fine architecture in which a coarse decoder predicts expression and pose trajectories, and a refinement module adds audio-conditioned stochastic expression residuals (z_t^r = z_t^c + β_t + γ_t ⊙ ε_t) while leaving the coarse pose parameters unchanged by construction (r̂_t = r̃_t).
- Empirical evaluation across two conversation benchmarks. Quantitative comparison, component ablations, and analyses of delay sensitivity, gate behavior, and blink dynamics on ViCo (D_test and D_ood splits) and L2L, using 1.50M parameters.
- Embodied evaluation. A robot-specific inverse-kinematic mapping plus relative motion calibration to deploy generated motion on an Ameca humanoid, supported by a Mean Opinion Score user study with N = 25 evaluators.
Main Findings
- Point-wise accuracy: REALM attains the lowest L1 expression and pose errors among evaluated methods on both datasets. On ViCo D_test it records L1 of 11.58 (expression) and 6.52 (pose); on ViCo D_ood, 13.76 and 5.02; on L2L, 9.67 and 2.41.
- Dynamic realism: On ViCo D_test, the inter-frame expression metric FID_Δfm drops from 8.71 (ListenFormer) to 3.91 with REALM. On L2L, REALM records 6.50 versus 13.41 for ListenFormer.
- Static distribution trade-offs: REALM achieves expression FD of 0.56 versus 0.57 for ListenFormer on ViCo D_test, and 1.47 versus a baseline range of 1.72–5.71 on D_ood. For rigid pose on D_test, however, RLHG retains a lower pose FD (0.72 vs 1.01).
- Static-vs-dynamic trade-off in a baseline: UniLS shows high spatial tracking error (L1 of 25.50 on ViCo D_test) yet strong dynamic smoothness (FID_Δfm of 4.48), illustrating that static coordinate fidelity and dynamic continuity do not move together.
- Speaker–listener correlation: REALM yields the lowest rPCC values across both benchmarks, though on L2L ListenFormer is comparable (0.008 / 0.010 versus 0.007 / 0.008).
- Motion variability: REALM closely matches ground-truth variance on ViCo D_ood and L2L. On ViCo D_test, ListenFormer's expression variance (0.142) is slightly closer to the reference than REALM's (0.133).
- Ablation — gating vs. shifted attention: Shifted Attention alone aligns reaction timing (reducing pose rPCC from 0.026 to 0.018) but omitting Gated Fusion lets the model over-react to acoustic energy, raising expression FD to 0.68. Combining both stabilizes the base trajectory, giving the lowest pose errors (L1 6.52, FD 1.01).
- Ablation — refinement module: RM improves expression FD to 0.62 in isolation and 0.56 in the full model, and sharply reduces expression FID_Δfm from ~12.50 to 4.11 on its own and 3.91 when integrated. Because RM is confined to the non-rigid subspace, pose kinematics are unaffected.
- Delay sensitivity: At 30 FPS (1 frame ≈ 33.3 ms), an 8-frame shift (≈267 ms) gives the best L1, FD and rPCC among tested settings {0, 4, 8, 12} frames. Increasing to 12 frames (≈400 ms) worsens results, particularly rPCC. The authors emphasize τ is the center of a soft prior, not a prescribed response latency.
- Gate behavior aligns with motion onsets: Across the ViCo test set, inference-time gate peaks versus ground-truth reaction onsets show 515 total reactions, a mean offset of -0.50 frames (-16.6 ms) and a median offset of 0.00 frames. Comparing top 10% versus bottom 10% gate frames within clips yields a mean velocity difference ΔV of 0.0051 with p < 0.001 (3.2 × 10^-5).
- Physical embodiment: Baselines such as ListenFormer frequently show deterministic smoothing or inappropriate mirroring, lowering Contextual Appropriateness and Naturalness. REALM records the highest Overall Preference in the user study.
- User study scores (1–5 scale): REALM achieves Naturalness 4.4, Audio-Visual Synchrony 4.2, Contextual Appropriateness 4.2, Overall Preference 4.0, versus ListenFormer at 3.6 / 3.8 / 4.0 / 3.6 and UniLS at 3.2 / 3.6 / 3.8 / 3.4. Pairwise score differences between REALM and the strongest baseline were statistically significant (p < 0.01, Wilcoxon signed-rank test).
Methodology in Plain English
The problem is framed as predicting a listener's motion at frame t given a window of speaker audio and the listener's preceding motion history over W frames. Each motion frame contains a non-rigid expression component and a rigid head-pose component.
The first design choice tackles timing. Rather than forcing a listener's response to line up with the speaker audio at the same moment, the authors add a soft attention bias that prefers speaker representations near a nominal lag τ behind the current frame. This is implemented as a shifted variant of ALiBi. Because it is a soft bias rather than a hard rule, the learned content-based part of attention can still select other admissible time steps when the conversation supports it. Future-indexed speaker frames are masked out, but the authors explicitly state the formulation is window-conditioned generation and does not by itself guarantee end-to-end streaming.
The second design choice tackles how strongly the speaker should matter. A learned gate produces a value between 0 and 1 that blends encoded listener history with the aligned speaker context. Low gate values lean on the listener's own momentum; high values lean on the speaker. The gate is trained jointly with the motion predictor, not supervised as a reaction detector.
The third design choice tackles what to generate. A coarse decoder produces a full base trajectory. A refinement module then samples a stochastic latent conditioned on the speaker context (with a learned mean shift and noise scale) and decodes it into an expression-only residual, which is added to the coarse expression while the coarse pose is passed through untouched. The two stages are optimized sequentially: coarse first for stable macro-motion, then refinement for local variation. Losses include reconstruction, temporal regularization, and adversarial terms.
Finally, for physical deployment the generated facial motion is mapped into robot actuator space through a fixed inverse-kinematic mapping with actuator-range clipping and resolution of overlapping control semantics, then smoothed and expressed as a displacement from a mechanical neutral pose. This preserves relative motion rather than absolute position, because human and robot neutral configurations differ.
Why This Matters
Listener behavior is a core non-verbal channel in conversation, and most generative avatar work targets the speaker. REALM's framing — that responses lag their cues, that speaker activity and listener motion do not vary proportionally, and that local facial events are stochastic — is a useful decomposition for anyone building dyadic interaction models. Its emphasis on physically grounding generated facial coefficients on real hardware matters because latent-space metrics do not measure physical plausibility.
Real-world applications implied by the paper's framing:
- Human-robot interaction: giving service or companion robots responsive, natural-looking listening faces, which the authors demonstrate on Ameca.
- Telepresence: animating a remote participant's avatar so it appears to be actively attending rather than frozen.
- Embodied conversational agents and digital avatars: making virtual agents signal comprehension and empathy through nods, eye contact, and expression shifts.
- Affective interaction research: providing a controllable generator for studying how listener responses shape a speaker's ongoing narration, a dynamic the paper cites from prior work on listener feedback.
Industry relevance: the framework uses only 1.50M parameters, retargets through a deterministic mapping rather than unconstrained end-to-end neural control, and the authors publish code at github.com/lipzh5/REALM with a demo video. The parameter count and the emphasis on mechanical safety bounds make the approach more plausible for embedded robot platforms than large end-to-end neural controllers, though the retargeting pipeline is described as robot-specific.
Future Directions
- Streaming interaction. REALM currently operates on offline conversational windows. The authors state that extending it to streaming human-robot interaction requires explicit control of temporal information access and end-to-end latency, and note they do not infer a streaming guarantee from the attention mask alone.
- Context-dependent delay priors. The nominal delay τ is a shared alignment prior that may not capture variation across individuals; developing context-dependent delay priors is named as a promising direction.
- Cross-embodiment retargeting. Physical deployment relies on a robot-specific retargeting pipeline; extending to different embodiments would require morphology-aware mappings and validation of actuator limits on each platform.
- Verifying stochastic capacity empirically. The paper leaves open whether the learned model actually uses its stochastic capacity to recover plausible local dynamics — this is stated as something evaluated empirically through motion and blink analyses rather than guaranteed by the parameterization. Note that the provided paper content is truncated mid-way through Appendix A, so the full appendices (B through F, covering losses, hyperparameters, mapping details, blink analysis, user-study procedure, and a detailed comparison of modeling choices) are not available in the supplied text and their contents are not summarized here.
Target Audience
Researchers and graduate students in human-robot interaction, embodied conversational agents, affective computing, and audio-driven facial animation; robotics engineers working on expressive humanoid faces who need a retargeting and perceptual-evaluation template; and practitioners who want a compact, deployable listener-motion generator rather than a large end-to-end model. Readers without a background in attention mechanisms, latent-variable models, or motion-capture coefficient representations will find the method sections demanding.
Authors’ abstract
Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the overall motion trajectory. Listener responses may follow preceding cues with a temporal lag, while brief expressions and blinks introduce variation that is difficult to predict deterministically. These challenges motivate a framework that combines history-aware temporal alignment with stochastic expression refinement. We propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework for audio-driven reactive listening. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio through a delay-centered attention prior and adaptive gating. A coarse decoder predicts a base motion trajectory, which is augmented by audio-conditioned stochastic residuals in the expression subspace while retaining the coarse pose parameters. Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics. Additional analyses examine delay sensitivity, gate behavior, and blink dynamics. Finally, deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment. Code: https://github.com/lipzh5/REALM Demo: https://youtu.be/Tf5mpd5S8VQ