Research
EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold
EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold Overview Research area: Computer vision, specifically interactive 3D head/talking-head motion generation with test-ti

- arXiv
- 2609.35616
- Published
- 2026-09-28
- Authors
- Junjie Chen, Fei Wang, Kun Li, Yiqi Nie, Xun Yang, Yanbin Hao, Linfeng Zhang, Meng Wang
AI summary
EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations UnfoldOverview
Research area: Computer vision, specifically interactive 3D head/talking-head motion generation with test-time training (TTT), plus a new benchmark for dyadic (two-person) conversation data.
Technical level: Advanced. The paper assumes familiarity with FLAME 3D head parameters, variational codecs, flow matching, transformer attention, and test-time training.
Scope in one sentence: The paper proposes a causal 3D head-motion generator whose internal parameters keep adapting during a live conversation using only unlabeled user video and both speakers' audio, together with a 455.95-hour benchmark for evaluating speaking and listening behavior across distribution shifts.
What This Paper Is About
Avatars that talk with people need to both speak and listen, and their motion should respond to how a specific conversation unfolds. Existing systems feed conversational context (avatar audio, both speakers' audio, or estimated user motion) into a model whose parameters are frozen at deployment, so the model never actually learns from the interaction. EvolvingAvatar instead learns a self-supervised signal from the conversation itself at test time and carries what it learns forward, without ever needing ground-truth user motion.
Key Contributions
- EvolvingAvatar, a causal generator that adapts at test time without motion labels through a Dyadic Context Prediction objective, combining persistent conversational adaptation (fast weights kept across intervals) with transient jaw adaptation (discarded after one interval).
- A region-structured causal FLAME codec that factorizes head motion into expression, neck pose, jaw pose, and a shared coordination latent, with codec-defined routing that sends expression and neck through a main flow-matching head and jaw through a separate jaw-specific head.
- InterHead-Bench, a unified 455.95-hour benchmark built by a framework called
dialog3d-factorythat merges single-view (mixed audio) and dual-view (separate audio) conversation recordings into aligned multimodal annotations, with 4,365 source participant IDs across Train, Dev, ID, OOD, and OOD-Hard splits. - An evaluation protocol combining parameter-space metrics (MSE, FD, P-FD, rPCC, SID, PDD, JDD), mesh-space metrics (LVE, MHD, FDD), and a blind forced-choice A/B perceptual study of speaking and listening behavior.
Main Findings
- Distributional statistics beat all baselines across splits. Expression and neck FD are lowest in both the speaking and listening states on every split, and on OOD-Hard jaw FD also leads. Speaking/listening FDD on OOD-Hard reaches 19.09/18.13 versus 20.91/19.84 for ARTalk. Findings for the all-frame setting are deferred to Appendix C.
- Performance improves as a conversation progresses. Across five equal-duration intervals, expression P-FD falls by 2.7% on OOD and 7.1% on OOD-Hard from the first interval to the last, while all four baselines worsen on OOD. The abstract separately reports a reduction of up to 11.1% in mismatch with recorded user–avatar expression statistics from the first interval on the hardest out-of-distribution split.
- Matched TTT controls support continued learning. Against "No write" (updates disabled), full TTT lowers ID expression P-FD by 20.4%; against "Freeze-1" (persistent updates stop after the first interval) by 4.2%; against "No carry" (fast state cleared after each group) by 8.6%. Neck P-FD instead increases by 8.9% against No write, which the authors attribute to DCP not directly constraining each region's motion statistics.
- Human raters prefer the method under shift. Overall-realism preference over UniLS rises from 58% on ID to 66% on OOD and 72% on OOD-Hard. On OOD-Hard, preference reaches 90% against DiffPoseTalk and ARTalk and 86% against DualTalk. The study used 10 raters across 60 trials with 95% finite-sample uncertainty intervals; the authors note the UniLS trend is descriptive because all three uncertainty intervals include equal preference.
- Point-error and perceptual quality diverge. DualTalk has lower MSE and LVE/MHD but lower motion coverage (SID), which the authors interpret as measuring agreement with one recorded response even though an interaction can admit several plausible responses. UniLS retains lower FDD on ID/OOD, which the authors describe as a remaining amplitude gap.
- Component ablations show complementary gains. With jaw flow disabled in both codecs, regional coding lowers ID expression and neck P-FD by 1.9% and 0.8% relative to unified coding, while jaw P-FD stays slightly higher. Jaw flow matching is associated with 5.1% lower jaw P-FD. Activity conditioning lowers all six FD/P-FD scores, including expression P-FD from 17.10 to 16.79. Jaw FM, activity, and RGB comparisons are described as exploratory across revisions.
- Model size is comparatively small. EvolvingAvatar reports 159.57M total and 54.82M trainable parameters, versus DualTalk at 647.27M/638.85M, UniLS at 422.88M/27.38M, ARTalk at 382.87M/37.89M, and DiffPoseTalk at 129.32M/110.55M.
- Visual context helps selectively. Appendix D reports that RGB improves expression and neck statistics over audio alone on OOD-Hard, but worsens expression statistics on ID/OOD.
Methodology in Plain English
Head motion is represented as a 106-dimensional FLAME vector per frame: expression coefficients (100 dimensions), neck rotation (3), and jaw rotation (3), excluding FLAME shape, global root pose, and eye pose. Time is cut into fixed-duration intervals, and the model emits motion at each interval boundary using only observations up to that point, so it never sees the future.
A frozen codec first turns ground-truth motion into a compact latent split into expression, neck, jaw, and shared coordination parts, trained with a region-balanced reconstruction loss plus a KL term whose average information rate is held under a budget. The generator then predicts these latents from user face frames and both speakers' audio, which are tokenized and fused through alternating cross-attention to audio and self-attention across N blocks.
The adaptation mechanism is the distinctive part. Inside the network, a small set of "fast weights" is updated before each use by an inner objective that predicts part of the observed context from another part (keys predicting values). The difference between the adapted and the original outputs is added back to the visual tokens. Persistent fast weights carry across intervals to capture conversational regularities; a separate, ungated transient branch reads the current audio and video tokens to adapt the jaw and then throws its fast weights away. A small activity encoder predicts speaking versus listening probabilities, and a gate derived from those probabilities scales only the added adaptation term, so the raw context is always available unchanged. Two flow-matching heads generate the latents, which the frozen decoder turns back into motion.
For the benchmark, dialog3d-factory standardizes and tracks faces, recovers each participant's audio from either separate dual-view tracks or a single-view mixture, annotates FLAME motion, transcripts, speaking states, interaction states, and turn statistics, and then validates and assigns segments to splits. ID holds out samples while allowing overlap with Train in participants or interactions; OOD holds out both; OOD-Hard tests transfer to a different recording domain (single-view RealTalk recordings, versus dual-view Seamless Interaction recordings).
Why This Matters
Research impact: The paper reframes interactive head generation from conditioning a fixed model on context to learning from context during generation. It also shows that a self-supervised objective over audiovisual context can substitute for target motion labels, which matters because user motion labels are typically unavailable in a live setting. The paired benchmark and matched ablations provide a way to separate reference-agreement metrics from perceived conversational quality.
Real-world applications:
- Virtual tutors that visibly respond to a student's behavior during a lesson.
- Avatars that support mental health conversations alongside human clinicians, which the authors flag as requiring clinician-guided study before deployment.
- Animation of avatars for dialogue systems, including full-duplex models that listen and speak simultaneously and multimodal models that combine audiovisual perception with streaming generation.
- Robotic facial actuation and talking-head synthesis driven by the generated 3D motion.
Industry relevance: Telepresence, gaming, virtual assistants, customer-service agents, and embodied robotics all need avatars whose nonverbal behavior stays synchronized with a specific interlocutor rather than replaying a generic style. A 159.57M-parameter model is also far smaller than the DualTalk baseline, which is relevant for deployment. The benchmark matters for the field: it defines a shared protocol on 455.95 hours of data for comparing speaking and listening under distribution shift.
Future Directions
- Studying continual learning in live interactions, where the model can observe how a user reacts to motion it generated. The current data cannot capture that response, since all recordings are of pre-existing human conversations.
- Deciding what to keep and when to revise it: distinguishing stable expressive habits from temporary emotional responses within a session.
- Better control over how visual context guides generation, given that RGB input helps on OOD-Hard but hurts expression statistics on ID/OOD.
- Closing the amplitude gap against UniLS on FDD for ID/OOD, and evaluating whether adaptation produces nonverbal responses appropriate to users' changing needs through clinician-guided studies that also track perceived support over longer conversations.
Target Audience
Researchers and graduate students in 3D head and talking-head generation, multimodal conversation modeling, and test-time or continual adaptation; HCI and affective-computing researchers studying nonverbal behavior in dialogue; and practitioners building conversational avatars, telepresence systems, or socially expressive robots who need a benchmark and a compact adaptive generator to compare against.
Authors’ abstract
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAvatar, a causal generator that uses test-time training to adapt to user face video and dyadic audio during interaction. Its dyadic context prediction objective provides a self-supervised learning signal from audiovisual context without target motion labels at test time. Persistent fast weights accumulate these updates within each conversation to guide motion generation, while transient jaw adaptation responds to current audiovisual context. Predicted speech activity controls how persistent adaptation guides motion. We also introduce InterHead-Bench, a unified 455.95-hour benchmark built from single-view and dual-view conversation videos. Experiments show improved conversational motion statistics over strong baselines. On the hardest out-of-distribution split, generation improves as conversations unfold, reducing mismatch with recorded user-avatar expression statistics by up to 11.1% from the first interval.