Skip to content
AI.info

Research

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation Overview Research area: Natural Language Processing / speech-to-speech dialogue modeling, specifically end-to-

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
arXiv
2609.36903
Published
2026-09-29
Authors
Ke Wang, Houxing Ren, Zimu Lu, Yunqiao Yang, Zhuofan Zong, Mingjie Zhan, Hongsheng Li

AI summary

MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation

Overview

  • Research area: Natural Language Processing / speech-to-speech dialogue modeling, specifically end-to-end full-duplex speech models (the Moshi paradigm), multi-party spoken conversation, and speech evaluation benchmarks.
  • Technical level: Advanced. The paper assumes familiarity with codec-frame-level modeling, parallel audio streams, forced alignment, and LLM-as-judge evaluation.
  • Scope: The paper releases a data engine plus a 57.6k-hour synthetic bilingual full-duplex corpus, a new benchmark (MultiTalkBench) for long multi-party bilingual dialogue, and a fine-tuned Moshi-style model (Moshi-MTB) trained on that corpus, evaluated against four open-source baselines.

What This Paper Is About

End-to-end full-duplex speech models like Moshi let a single system listen and speak at the same time, but the paper argues they fall short of real deployment in two entangled ways: they do not stay coherent over long conversations, and they cannot handle more than two speakers. The authors note that realistic settings — meetings, group lessons, family dinners, social-robot reception — are inherently long-horizon and multi-party at once, yet open multi-party conversational speech corpora total only a few hundred hours and existing benchmarks are either passive-listening long-audio tests or short dyadic speech-to-speech tests.

The goal is therefore to extend the Moshi paradigm along the long-horizon and multi-party axes simultaneously in both English and Chinese, by supplying all three missing pieces at once: training data, an evaluation benchmark, and a trained model. The authors state that open multi-party corpora such as AMI, ICSI, AISHELL-4, and AliMeeting together amount to only a few hundred hours and were not designed for codec-frame-level full-duplex modeling, and that every recent Moshi-style system instead falls back on proprietary in-house data or non-redistributed TTS-synthesized stereo dialogue.

Key Contributions

  1. An automatic data engine and an open 57.6k-hour synthetic training corpus for long, multi-party, English–Chinese full-duplex dialogue. The engine produces parallel-stream audio with controllable length, participant count, conversational dynamics (turn-taking, overlap, backchannels, interruption, addressee shifts, long-range co-reference), and language. The authors state this exceeds all prior open multi-party conversational speech corpora by more than an order of magnitude.

  2. MultiTalkBench, described as the first benchmark to jointly evaluate long, multi-party, and bilingual full-duplex dialogue, built from real human recordings. It tests interactive conversations with explicit probes for long-range entity tracking and topic coherence, one-model-many-user multi-party interaction with quantitative addressee-selection and turn-taking metrics, and English–Chinese bilingual ability.

  3. A bilingual Moshi-style full-duplex model (Moshi-MTB) trained with a two-phase recipe on the released corpus, which the authors report substantially outperforms open-source baselines including Moshi (Moshiko-7B), MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct.

  4. A human-evaluation validation of the LLM-as-judge protocol, comparing two LLM judges against human raters on 100 sampled model outputs to show that automated scoring tracks human judgment.

Main Findings

  • Overall gains over the base checkpoint: Training the off-the-shelf Moshiko-7B checkpoint on MultiTalkPT + MultiTalkFT raises the overall MultiTalkBench Final Score by +182.2%, from 4.66 (Moshiko-7B) to 13.15 (Moshi-MTB).

  • It beats all four open-source baselines on the aggregate score: Moshi-MTB scores 13.15, versus MiniCPM-o-4.5 at 11.05, Qwen3-Omni-30B at 4.94, and PersonaPlex-7B at 2.52.

  • Multi-party competence improves sharply: Multi-Party IQ rises from 2.41 (Moshiko-7B) to 8.56, a +255.2% change, narrowly trailing MiniCPM-o-4.5 at 9.00. Moshi-MTB leads MiniCPM-o-4.5 on Speaker Tracking (14.56 vs. 13.60), Group Awareness (15.80 vs. 14.40), and Noise Resistance (12.77 vs. 8.23).

  • Role-conditional behavior improves: The Role-conditional group rises from 6.62 to 12.16 (+83.7%), exceeding MiniCPM-o-4.5 (7.81) by +55.7%.

  • The largest single gain is on Chinese: The ZH final score rises from 0.36 on Moshiko-7B to 16.27 on Moshi-MTB, a factor of about 45×, while English is preserved (EN final from 7.82 to 10.87, +39.0%).

  • The largest model in the comparison does poorly on multi-party behavior: Qwen3-Omni-30B reaches only 2.81 on Multi-Party IQ and 5.08 on Role-conditional, which the authors read as evidence that multi-party behavior is bounded by training distribution rather than parameter count.

  • Participation scores tell a different story: PersonaPlex-7B attains the highest Participation score (97.87), yet its General, Multi-Party IQ, and Role-conditional scores all fall in the bottom quartile, yielding the lowest final score of the compared models (2.52).

  • Large headroom remains: All five compared models stay far below the human reference of 68.74; Moshi-MTB reaches 13.15.

  • More multi-party fine-tuning data helps monotonically: With the earlier stage fixed, increasing the MultiTalkFT fraction from 0% to 25%, 50%, 75%, and 100% raises the Final Score from 7.95 to 10.39, 11.78, 12.77, and 13.15.

  • Data diversity matters at matched scale: In size-matched 800-hour comparisons, the score is 10.39 with a random subset, 8.30 with only three-speaker conversations, and 9.55 with only the shortest conversations.

  • Performance degrades with group size and length: Moshi-MTB scores 13.44 on three-speaker samples versus 11.96 on larger groups, and 21.23 on conversations under 30 minutes versus 12.07 on longer ones.

  • The benefit transfers across architectures: F-Actor's Final Score increases from 6.94 to 15.16 after training with MultiTalk.

  • Interaction timing improves: Moshi-MTB achieves 870 ms response onset latency, 2230 ms stop latency, 94.1% backchannel continuation, and 98.8% side-conversation ignoring, versus 782 ms, 5475 ms, 93.7%, and 92.0% for Moshiko-7B.

  • LLM judges track human raters closely: The Gemma-4-31B judge achieves Spearman ρ ≥ 0.84 on every metric group, peaking at ρ = 0.898 on the Final Score; inter-judge agreement stays above ρ = 0.72 across all groups, reaching ρ = 0.847 on the Final Score.

Methodology in Plain English

Building training data. The authors build an automatic pipeline that starts from four pools of seed material: 433K emotion-labeled sentences from dair-ai/emotion, 1.7M question–answer pairs from AM-DeepSeek-R1-0528-Distilled and AM-Qwen3-Distilled, 32K character profiles normalized by prompting gemini-2.5-pro, and 840 discussion topics across 84 industries. From these, an LLM makes two separate calls: the first designs the scenario, cast, and interaction trajectory; the second writes the full dialogue under that fixed world model. Conversations have between 2 and 9 participants (most commonly 3), 50 to 300 turns, and are in English or Chinese. The prompt enforces a spoken-length mix of 60% short turns (1–15 words), 30% medium (15–40), and 10% long, plus an interaction protocol for backchannels, interruptions, overlap, and explicit silence turns. A four-layer quality filter rejects bad scripts with structured reason codes.

Rendering audio. Each turn is synthesized with IndexTTS2 using gender-matched voice prompts drawn without replacement so no two participants share a voice, and an eight-dimensional emotion vector. The per-utterance waveforms are assembled into multi-channel tracks with the assistant on channel 0; each speaker's turns are placed sequentially without overlapping their own channel, while cross-speaker overlap is freely allowed. A gap-compression pass makes adjacent cross-speaker boundaries overlap by 0.2 to 0.6 seconds, or 1 to 2 seconds after an interruption. Word-level timestamps come from forced alignment with the Montreal Forced Aligner.

Training the model. Starting from the public kyutai/moshiko-pytorch-bf16 checkpoint, they run two phases without changing Moshi's dual-stream architecture. Phase one pre-trains on MultiTalkPT (54.4k hours) with the user stream also supervised, at roughly 60% English and 40% Chinese, for 5,000 steps with a peak learning rate of 3×10⁻⁵. Phase two fine-tunes on MultiTalkFT (3.2k hours of multi-party data) for 200 steps with a peak learning rate of 2×10⁻⁶ for the temporal transformer and 4×10⁻⁶ for the depth transformer, mixing all non-target speakers into the user channel. Both phases use a global batch size of 21 hours of audio, AdamW with a one-cycle schedule, and DNS-Challenge noise mixed into the user channel.

Evaluating. MultiTalkBench is built from real recordings drawn from AMI, ICSI, CHiME-6, AISHELL-4, AliMeeting, AISHELL-5, and MagicData-RAMC, keeping only sessions with per-speaker close-mic or lapel audio. From the filtered pool they generate 104 evaluation samples. The model plays a designated named participant, and its generated utterances are spliced into an ordered transcript scored by an LLM judge across four groups: General, Multi-Party IQ, Role-conditional, and a mechanical Participation score based on span-ratio participation matching. The three LLM-judge groups are weighted equally at 1/3 each.

Why This Matters

  • Impact on research: The paper argues the field has been blocked by proprietary in-house data, and that no prior work trained or evaluated a one-model, many-user, long-horizon end-to-end full-duplex system. It asserts that multi-party competence appears bounded by training distribution rather than parameter count, and it releases the data engine, corpus, and benchmark to make this axis testable. The reported 13.15 versus 68.74 human reference suggests the benchmark is far from saturated.

  • Real-world applications named or implied by the paper:

    • Accessibility: real-time captioning and turn-taking support for hearing-impaired users in group settings.
    • Education: automated facilitators in group lessons.
    • Assistive robotics: social robots handling reception, family, or care scenarios involving multiple humans.
    • Meetings: the benchmark is framed around multi-party meetings with designated roles (Facilitator, Driver, Collaborator, Evaluator) and requires addressee selection and cross-speaker information integration.
  • Industry relevance: The work targets deployed conversational agents that run continuously rather than over short clips, positions itself against commercial references such as GPT-4o and Gemini Live, and benchmarks against production-oriented open models including MiniCPM-o-4.5 and Qwen3-Omni-30B-A3B-Instruct. Lowering the data barrier is framed as the main practical unlock for academic and smaller-lab research.

Future Directions

  • Dialogue-aware synthesis for closer train–test acoustic match: The authors note that the training corpus is synthetic while the benchmark is real, making the evaluation a cross-domain assessment by construction, and suggest dialogue-aware TTS or real-prosody grafting as a natural next step.
  • Extension to additional languages: Reported results cover English and Chinese separately; the benchmark does not test intra-sentential code-switching, and the authors warn that extension to lower-resource languages may be limited by TTS quality, voice diversity, script fluency and cultural appropriateness, language-dependent codec and ASR errors, and the availability of real multi-party recordings and native-speaker judge validation.
  • Scaling the recipe to larger backbones: The authors state that their 7B-parameter backbone bounds the headroom data alone can recover, and that they have not evaluated the recipe on larger backbones.
  • Establishing a scaling law for the corpus and isolating synthetic-to-real variation: The results are said to support benefits from MultiTalkFT scale and diversity, cross-architecture transfer, and improved interruption handling, without establishing a scaling law for the entire corpus or isolating every source of synthetic-to-real variation.

Target Audience

Speech and dialogue-systems researchers working on full-duplex, streaming, or speech-to-speech models; teams building multi-party conversational agents, meeting assistants, or social robots; and benchmark or evaluation researchers interested in long-form, interactive, multilingual speech evaluation. It is also relevant to data-engineering groups looking for an open, controllable pipeline for generating codec-frame-aligned multi-speaker dialogue, and to anyone who needs to understand where open full-duplex models currently stand relative to human performance on multi-party conversation.

Authors’ abstract

End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{https://huggingface.co/datasets/MultiTalk/MultiTalkPT}{MultiTalkPT}$ and $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkFT}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkBench}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.

Read the original paper