Research
GLARE: Generating Listening Heads with Appropriate Reactions
Overview Research area: Computer vision / generative modeling for dyadic human conversation — specifically listening head generation, the task of animating the person who is listening (not speaking) i

- arXiv
- 2609.40317
- Published
- 2026-09-30
- Authors
- Zikai Liao, Yumin Suh, Yi Ouyang, Yi-Lun Lee, Yi-Hsuan Tsai, Zhaozheng Yin
AI summary
Overview
Research area: Computer vision / generative modeling for dyadic human conversation — specifically listening head generation, the task of animating the person who is listening (not speaking) in a two-person conversation.
Technical level: Advanced. The paper assumes familiarity with flow-matching transformers, latent motion autoencoders, diffusion/DiT-style architectures, audio LLM conditioning, and event-level detection metrics.
Scope (one sentence): The paper builds a large annotated listening-reaction dataset from RealTalk and Seamless Interaction, proposes an audio-driven listener generator (GLARE) with prosody conditioning and a temporal reaction loss, and introduces reaction-centric evaluation metrics that measure whether a listener reacts with the right type at the right time and with good visual quality.
What This Paper Is About
Most talking-head research animates the speaker, but in a real face-to-face conversation the listener's non-verbal feedback — nodding, smiling, laughing, frowning, head shaking, or looking surprised — is what makes the interaction feel natural. The problem is that existing dyadic conversation datasets do not label listener reactions at the event level (type, start time, end time), and existing evaluation metrics (FID, FVD, LPIPS, PSNR, SSIM) measure global visual realism rather than whether a reaction was appropriate and correctly timed. The goal of this paper is to close that gap on three fronts at once: reaction-aware data, reaction-aware modeling, and reaction-aware evaluation.
Key Contributions
-
A listening-head-specific dataset. Curated from RealTalk and Seamless Interaction into 107,149 speaker audio / listener video pairs — approximately 147 hours of paired speaker–listener video — with 64,557 event-level reaction annotations spanning six categories (nodding, head shaking, smiling, laughing, frowning, surprised).
-
GLARE, an audio-driven reacting listener. A latent flow-matching transformer (FMT) baseline with a dedicated prosody conditioning branch (prosody extracted with Qwen2-Audio-7B-Instruct) and a temporal reaction loss that explicitly supervises frame-wise reactions.
-
A reaction-oriented evaluation protocol. Four new metrics — R-F1 (reaction occurrence/type), R-tIoU (temporal alignment), R-ATD (asymmetric temporal deviation), and R-FID (visual quality inside reaction regions) — evaluated alongside conventional generation metrics.
-
Extensive comparison and ablations. Comparisons against L2L, DIM, ViCo, ListenFormer, and DyStream on both RealTalk and Seamless, plus ablations over the two proposed modules, the reaction loss coefficient, and per-reaction-category human evaluation. The code, annotation pipeline, and processed dataset are stated to be released at https://github.com/lzk901372/glare.
Main Findings
-
Reaction-level metrics improve over all compared baselines on RealTalk. GLARE reaches R-F1 0.594, R-tIoU 0.704, R-ATD 57.388, and R-FID 15.192, versus the strongest competing values from DyStream (R-F1 0.535, R-tIoU 0.694, R-ATD 63.171, R-FID 15.337).
-
Conventional visual metrics also improve on RealTalk, with one exception. GLARE achieves PSNR 17.972, FID 35.697, FVD 142.454, Var 2.916, LPIPS 0.454, rPCC 0.227, and DI-Sync 0.245 — best among the listed methods — while DyStream records a higher SSIM of 0.611 versus GLARE's 0.601.
-
Seamless is harder and shows the same pattern. GLARE reaches PSNR 14.245, SSIM 0.473, FID 27.288, FVD 193.281, Var 2.447, LPIPS 0.371, rPCC 0.301, DI-Sync 0.205, R-F1 0.460, R-tIoU 0.582, R-ATD 102.113, and R-FID 28.857. The paper notes DyStream obtains slightly higher Var (2.452), and ListenFormer reports a slightly higher DI-Sync (0.212).
-
Per-class reaction behavior is uneven. In Table 3 (compared against DyStream), laughing is described as easier to detect and align, while subtle motions such as nodding and head shaking are more sensitive to timing errors; surprised is reported as the most challenging category due to ambiguity and sparsity, and smiling benefits from its larger data scale. All Seamless scores are consistently lower than RealTalk scores.
-
Roughly equal reaction counts across categories. RealTalk per-class annotated counts are: nodding 3033, head shaking 3588, smiling 4804, laughing 2397, frowning 2745, surprised 1860. Seamless counts are: nodding 7197, head shaking 9202, smiling 10963, laughing 5923, frowning 7241, surprised 5604.
-
Ablation: the reaction loss does most of the work. Starting from neither module (PSNR 13.113, R-F1 0.404, R-tIoU 0.599, R-ATD 89.984, R-FID 23.706), adding only prosody improves visually (PSNR 14.071, FVD 159.745, Var 2.206) but worsens R-ATD to 100.342. Adding only the reaction loss gives large gains (PSNR 17.212, DI-Sync 0.236, R-F1 0.579, R-tIoU 0.695, R-ATD 55.657, R-FID 17.885). Using both gives the best overall configuration (PSNR 17.973, SSIM 0.601, FID 35.691, FVD 142.456, DI-Sync 0.245, R-F1 0.594, R-tIoU 0.704, R-ATD 57.386, R-FID 15.192).
-
λ_react = 0.05 is the reported optimum. At 0.01 the model gets PSNR 15.446 and R-F1 0.572; at 0.05 PSNR 17.973 and R-F1 0.594; at 0.1 R-F1 0.613 with R-ATD 64.790; at 0.2 performance degrades (R-ATD 69.131); at 0.5 and 1.0 both visual quality and reaction metrics degrade further.
-
Prosody conditioning raises reaction occurrence but not timing precision. Table 6 reports R-F1 gains across all six reaction types with prosody on, largest for laughing (+29.31%), while most categories show degraded R-tIoU and higher (worse) R-ATD. The paper attributes this to prosody reflecting the speaker's affective state rather than dictating when a listener should respond.
-
Human evaluation is favorable. Across 100 generated RealTalk test videos (15.94 s average) containing 165 reaction events, 23 participants produced 165 × 23 = 3,795 reaction-participant annotations. Overall Human Agreement Rates are 0.899 for reaction naturalness, 0.894 for contextual appropriateness, and 0.935 for timing plausibility. Highest per-category naturalness is head shaking (0.969) and lowest is surprised (0.824); contextual appropriateness is highest for nodding (0.975) and lowest for laughing (0.836).
Methodology in Plain English
Data curation. The authors start from RealTalk and Seamless Interaction. They filter out bad video (scene transitions, black frames, multiple people, occlusions, unstable face crops), crop each participant into a portrait-centered video, and run a manual face-quality check. They then figure out who is speaking when: for RealTalk they perform audio separation, denoising, and speaker diarization; for Seamless they use the provided audio streams and diarization with timestamp correction. Each conversation is broken into speaker-dominant intervals, and the active speaker's audio is paired with the synchronized portrait video of the other person, who is treated as the listener.
Reaction annotation. A visual reaction detector is applied to cropped listener clips and outputs a dense per-frame score vector with six channels, one per reaction category, combining head-motion cues, facial landmark dynamics, facial expression cues, and temporal smoothing. Continuous high-confidence regions become events with a type, a start time, and an end time; overlaps between classes are resolved by keeping the dominant reaction. Manual verification removes unreliable detections.
Model. The listener reference image is encoded by a frozen LIA autoencoder into a motion latent and an identity latent; the speaker audio is encoded by Wav2Vec2 into audio latents and by an emotion encoder into emotion latents. A prosody branch has Qwen2-Audio-7B-Instruct analyze the audio and produce timestamps of prosodic fluctuation, converted into a frame-aligned scalar in [0, 1] and projected through an MLP and Sigmoid into a prosody latent. These four latents are channel-concatenated into the driving condition. A DiT-style transformer predicts conditional vector fields, the fields are integrated into a listener motion latent, and the latent is decoded into video frames.
Temporal reaction loss. Hidden latents for the current frames are mapped through an MLP to a 192-dimensional representation, split into six class-specific subspaces of 32 channels each, and each reduced to a per-frame reaction probability in [0, 1] with a class-specific linear layer and sigmoid. These predictions are supervised against the annotated frame-wise reaction targets with a Smooth-L1 loss, so reaction structure is baked into the hidden motion representation rather than only into pixels.
Training and metrics. The total objective combines the flow-matching loss, a temporal velocity consistency regularizer, and the reaction loss. Training uses mixed precision with Accelerate on 4× L40s GPUs, AdamW at lr = 5e-4, gradient clipping at 1.0, cosine decay with warmup, batch size 256, and 650 epochs, at 25 FPS, 16 kHz audio, 512×512 inputs normalized to [-1, 1], with a 9:1 train/test split. Evaluation adds RI-F1, R-tIoU, R-ATD, and R-FID to standard metrics; event matching requires identical reaction type and temporal IoU above τ = 0.5, R-ATD penalizes early/truncated reactions more heavily than late ones via asymmetric weights α > β, and R-FID computes the Fréchet distance only on frames inside reaction intervals.
Why This Matters
Impact on research. The paper reframes listening head generation as a reaction-modeling problem rather than a pure video-synthesis problem, and supplies the three ingredients — event-level data, a reaction-supervised objective, and reaction-centric metrics — that previous work lacked. Its explicit acknowledgment that listener feedback is "multi-valid" (a plausible response need not match one reference timestamp) points toward evaluation designs that go beyond naive reconstruction.
Real-world applications:
- Virtual agents and avatars in video calls, telepresence, and customer-service chatbots that must appear to be listening attentively.
- Dubbing and localization pipelines that need listener reactions to remain consistent with re-recorded or translated speaker audio.
- Remote meeting and interview tooling that analyzes or synthesizes listener feedback for training and coaching.
- Social-skills and clinical training simulations where appropriate non-verbal feedback timing matters (for example, practicing conversation with autism-spectrum learners or interview candidates).
Industry relevance. Any product generating synthetic humans in conversation — game NPCs, VR/AR avatars, synthetic media, educational tutors, and conversational AI with a visual face — needs listeners that react on cue. The metrics proposed here give teams a way to benchmark that behavior rather than relying on global visual fidelity, and the released dataset and annotation pipeline lower the cost of building reaction-aware systems.
Future Directions
- Distinguishing prosody from semantics. Prosody conditioning improves R-F1 across all six categories but often degrades R-tIoU and R-ATD, which the authors attribute to prosody indicating affect rather than whether or when a listener should react. Adding semantic or contextual cues is a natural next step.
- Handling the hard categories. Surprised remains the most difficult class (lowest naturalness HAR at 0.824 and the smallest RealTalk count at 1860 annotations), and laughing has the lowest contextual appropriateness (0.836). Better modeling of sparse, ambiguous, and humor-dependent reactions is open.
- Improving the Seamless subset. Every reported score is lower on Seamless than RealTalk, and GLARE trails ListenFormer on Seamless DI-Sync (0.205 vs 0.212) and DyStream on Seamless Var (2.447 vs 2.452) — closing that gap is an explicit target.
- Beyond the ground truth. Since a single reference reaction is only one of several plausible responses, extending evaluation toward multi-valid or distributional reaction assessment — rather than comparison against one ground-truth annotation — is left as future work.
Target Audience
Researchers and engineers working on talking-head and listening-head generation, dyadic and multi-modal conversational AI, and human behavior analysis; dataset builders interested in event-level annotation pipelines for non-verbal behavior; and applied teams building avatars, virtual agents, or conversational interfaces that need believable listener feedback. Readers without a background in flow matching or audio-language models will need to consult the cited prior work (FLOAT, LIA, Wav2Vec2, Qwen2-Audio) to follow the methodology section closely.
Authors’ abstract
While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker-listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.