Research
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
Overview Research area: Spoken (full-duplex) dialogue modeling, multimodal machine learning, audio-visual speech recognition, turn-taking prediction. Technical level: Advanced. The paper assumes famil

- arXiv
- 2511.11124
- Published
- 2025-11-14
- Authors
- Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota
AI summary
Overview
- Research area: Spoken (full-duplex) dialogue modeling, multimodal machine learning, audio-visual speech recognition, turn-taking prediction.
- Technical level: Advanced. The paper assumes familiarity with streaming speech models, audio codecs as tokenizers, multimodal LLM fine-tuning, and turn-taking metrics.
- Scope: The paper introduces AV-Dialog, a streaming audio-visual dialogue framework built on Llama3-8B that uses lip-centric visual cues plus acoustic tokens to track a target speaker, predict turn boundaries, and generate responses in noisy multi-speaker conditions.
What This Paper Is About
Current spoken dialogue models listen but do not look, so they lose track of the target speaker, respond to the wrong person, and misjudge when to take a turn in noisy, overlapping-speech settings. AV-Dialog adds a visual stream (lip movements from first-person video) on top of audio so the model can isolate the intended speaker and decide when to speak. The goal is a streaming dialogue system that transcribes, predicts turn-taking, and generates coherent responses under the same "cocktail party" conditions where humans combine hearing and sight.
Key Contributions
- First audio-visual spoken dialogue framework. AV-Dialog jointly performs streaming speech recognition, turn-boundary prediction, and response generation, processing audio and video in 40 ms chunks. It is presented in two variants: a dual architecture (an AV understanding module plus a separate text backbone) and a unified architecture (a single model doing AV understanding, turn prediction, and generation).
- Acoustic tokenization instead of semantic tokenization. Rather than speaker-invariant semantic tokens such as HuBERT, AV-Dialog uses general-purpose acoustic tokens from the Descript Audio Codec (DAC), encoding each 40 ms chunk into 16 codebooks. The authors argue this preserves voice characteristics and enables inherent speaker differentiation under interference.
- Multi-task, multi-stage training recipe with synthetic mixing. Stage 1 aligns text, audio, and vision through four objectives (text continuation, speech comprehension/ASR, audio captioning, AVSR). Stage 2 fine-tunes on real audio-only and audio-visual conversations for turn-taking and dialogue. Synthetic mixing augmentation simulates noisy multi-speaker environments.
- Explicit turn-event supervision. The model predicts a turn-event stream alongside transcription, using PairwiseTurnGPT's taxonomy (normal turn, overlapping turn, backchannel) with
<SOT>,<SOB>, and<EMP>tokens. The paper shows this supervision is essential for the dual setup and also improves the unified model.
Main Findings
- Visual input improves turn-taking under interference. Adding vision to the audio-only model improves turn-taking accuracy by 1.3% (Clean), 8.1% (BG), and 13% (Interf). The audio-visual model reaches a response ratio of 74.5% (Clean), 78.3% (BG), and 78.8% (Interf), versus roughly 50% for the Moshi baseline.
- Headline claim from the introduction: adding the visual modality boosts turn-taking prediction accuracy from 54% to 79% in the presence of interfering speakers.
- Streaming AVSR on VoxCeleb2 (WER %). Auto-AVSR (A+V): 26.8 Clean, 48.2 BG, 71.8 Interf. Ours (A+V): 17.4 Clean, 35.6 BG, 38.8 Interf. Ours (A): 18.0 / 60.2 / 76.3. Ours (V): 87.1 across all three conditions.
- Streaming AVSR on InterAct (WER %). Ours (A): 28.6 / 68.0 / 92.2. Ours (V): 67.8 across all conditions. Ours (A+V): 16.3 / 37.4 / 30.8.
- Acoustic tokens beat semantic tokens. With DinoSR semantic tokens, AVSR WER under interference reached 239.2 (audio-only) and 67.0 (audio-visual); DAC acoustic tokens gave 63.4 (audio-only) and 30.8 (audio-visual). The introduction states that replacing semantic with acoustic tokens reduces AVSR word error rate from 67% to 31.7% under strong multi-speaker interference. Acoustic tokens also improved response ratio (74.5% / 76.9% / 78.8% for A+V versus 69.5% / 49.1% / 47.8% for DinoSR A+V).
- Dialogue semantics (Table 4). Moshi pickup ratio: 23.8% / 19.4% / 18.4%; SE+Moshi: 26.4% / 24.5% / 19.1%. Ours (ICL): 66.6% / 68.1% / 67.8% — the highest. Ours (IT): 32.5% / 30.3% / 36.7%. Ours (Unified): 29.6% / 35.5% / 31.3%. Perplexity: ICL 25.8 / 24.6 / 23.1 and IT 23.2 / 24.0 / 23.8, versus Moshi 44.1 / 52.4 / 46.4.
- Human evaluation (N=18, ITU-T P.808, 5-point scale). N-MOS: SE+Moshi 2.39, Ours (Dual+ICL) 4.14, Ours (Unified) 3.54, GT 3.92. H-MOS: SE+Moshi 2.10, Ours (Dual+ICL) 4.09, Ours (Unified) 3.02, GT 3.62. The abstract reports this as a +1.75-point MOS improvement in naturalness and +1.99-point gain in relevance/helpfulness.
- Stage 1 training is critical. Removing Stage 1 raises AVSR WER to 58.9 / 95.1 / 86.7, versus 16.3 / 37.4 / 30.8 with it. Excluding the audio-only dialogue dataset in Stage 2 also hurts (22.6 / 37.8 / 31.8).
- Explicit turn-taking supervision helps the unified model. Response ratio with explicit turn-taking: 68.1 / 75.6 / 75.9; without it: 48.0 / 35.1 / 38.0. LLM-evaluator pickup ratio also drops (29.6 / 35.5 / 31.3 versus 29.0 / 22.1 / 18.2).
- Ground-truth reference points. The GT median floor-transfer offset (FTO) is around 1.5 s in the InterAct dataset. The response ratio metric counts FTOs within –2 s to 3 s, a range containing roughly 90% of FTOs in the InterAct test set.
- Algorithmic latency. AV-Dialog's DAC tokenizer and visual encoder operate causally at 25 Hz, but the AV-HuBERT visual encoder adds a 2-frame lookahead, giving roughly 120 ms algorithmic latency. Moshi processes 80 ms chunks with about 80 ms latency.
Methodology in Plain English
The system starts from a pretrained Llama3-8B text model and teaches it to handle three things at once: what the user said, when the user is done, and when the agent should speak.
For input, audio is broken into 40 ms chunks and encoded into 16 parallel streams using the DAC audio codec. On the video side, a face detector (dlib) finds face regions in first-person video, and a pretrained AV-HuBERT model extracts lip-centric features. Audio embeddings, visual embeddings, and the previous step's text and turn-token embeddings are projected into the same dimension and summed into a single input embedding at each timestep.
Training happens in two stages. Stage 1 mixes four tasks to align modalities: text continuation (to preserve language ability), ASR on monaural speech, audio captioning on general audio, and AVSR on a monadic audio-visual dataset. Stage 2 fine-tunes on real conversations — audio-only and audio-visual — for streaming AVSR and turn-taking prediction. To handle noise, samples are synthetically mixed: 20% stay clean, 40% get background noise, and 40% get 1–4 interfering speakers, with input SNR sampled uniformly between –8 dB and 8 dB.
For the dual model, the AV understanding module streams recognized text and turn tokens into a separate text LLM, which either uses in-context learning with few-shot dialogue examples or is instruction-tuned on human conversations. The unified model instead emits turn events and response tokens on the same stream, dropping the AVSR stream. Generated text is turned into speech with Moshi's Mimi streaming TTS.
Evaluation covers three conditions — Clean, BG (noise from MUSAN at –8 dB to 12 dB SNR), and Interf (1–4 interfering speakers at the same SNR range) — using word error rate, response ratio, FTO error, median FTO, perplexity, an LLM judge (Prometheus-7b-v2.0) for response pickup, and a human MOS study.
Why This Matters
- Research impact: The paper argues that visual grounding and acoustic (rather than semantic) tokenization are the missing pieces for speaker-aware dialogue in noisy settings. It also reports that a cascade (dual) architecture produces better responses than a unified one, matching earlier observations in speech-to-speech dialogue.
- Real-world applications:
- Video conferencing assistants that follow the correct speaker when several people talk at once.
- Smart glasses, hearables, or always-on wearables with a first-person camera that help users follow conversations in noisy rooms.
- Voice interfaces for smart speakers or robots that need to know when to speak and when to stay silent.
- Accessibility tools for hard-of-hearing users who need speaker-attributed transcription and responses.
- Industry relevance: The work combines an off-the-shelf codec (DAC), an off-the-shelf visual encoder (AV-HuBERT), an off-the-shelf LLM (Llama3-8B), and an off-the-shelf TTS (Mimi), so the recipe is largely portable to existing production stacks. The dual design also lets developers swap in a different text LLM or API without retraining the AV understanding module. The privacy discussion around always-on audio-visual capture is directly relevant to deployed products.
Future Directions
- Non-verbal cues. The model does not explicitly model non-verbal auditory signals (laughter, sighs) or visual signals beyond lip movement (facial expressions, gestures). Adding these could make interactions more human-like.
- Robust visual encoding. Poor lighting, occlusions such as hands covering the mouth, and extreme head poses impair lip-movement extraction. The authors call for lip encoders robust to these conditions.
- Lower algorithmic latency. AV-Dialog's roughly 120 ms latency is dominated by the visual encoder's 2-frame lookahead. Pretraining a visual encoder with a smaller or zero lookahead window could reduce this toward Moshi-like latency.
- Improving the unified model. The unified model trails the dual/ICL setup in response quality, which the authors attribute to limited real conversational data and the casual, low-quality nature of that data. Better data or training strategies for joint turn-taking and generation remain open.
Target Audience
Researchers and engineers working on spoken dialogue systems, full-duplex or streaming conversational agents, audio-visual speech recognition, and multimodal LLM fine-tuning. It is also useful for product teams building conferencing, wearable, or voice-assistant systems that must operate in noisy, multi-speaker environments, and for readers interested in how acoustic versus semantic speech tokens affect robustness.
Authors’ abstract
Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses. By combining acoustic tokenization with multi-task, multi-stage training on monadic, synthetic, and real audio-visual dialogue datasets, AV-Dialog achieves robust streaming transcription, semantically grounded turn-boundary detection and accurate responses, resulting in a natural conversational flow. Experiments show that AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction, and enhancing human-rated dialogue quality. These results highlight the power of seeing as well as hearing for speaker-aware interaction, paving the way for {spoken} dialogue agents that perform {robustly} in real-world, noisy environments.