Skip to content
AI.info

Research

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

Overview Research area: Spoken dialogue systems — specifically evaluation of full-duplex (listen-while-speaking) speech models in multi-party conversations (cs.SD; arXiv:2609.31948v1). Technical level

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
arXiv
2609.31948
Published
2026-09-25
Authors
Chengqian Ma, Wenhao Feng, Weixuan Jin, Gaole Dai, Tianyu Xie, Yuexiao Ma, Zhaolu Kang, Xiangyu Zhao, Xiawu Zheng, Fei Chao

AI summary

Overview

Research area: Spoken dialogue systems — specifically evaluation of full-duplex (listen-while-speaking) speech models in multi-party conversations (cs.SD; arXiv:2609.31948v1).

Technical level: Intermediate. The conceptual setup is accessible, but the metric definitions, duplex streaming protocol and VAD-based timing measurements assume familiarity with speech-model evaluation.

Scope: The paper introduces Duplex-MPE, a benchmark of 2,000 three-or-four-speaker scenarios (4,000 evaluated conversations) that tests whether a spoken assistant can decide when to answer, stay silent, and stop speaking in a shared multi-party conversation.

What This Paper Is About

Real-time full-duplex speech models can listen while they talk, so they are no longer bound to rigid turn boundaries. Existing benchmarks, however, mostly evaluate an assistant serving one designated user, not an assistant taking part in a conversation among several people where any participant might address it, someone else, or nobody in particular. This paper builds a benchmark that measures selective participation — when the assistant should speak, when it should hold silence, and when it should release the floor — using continuous audio with no transcripts, speaker labels or turn boundaries supplied to the model.

Key Contributions

  1. A benchmark for shared multi-party dialogue. Duplex-MPE provides 2,000 scenario pairs where the same request is rendered once with the assistant explicitly named and once with the addressee left to be inferred from context, so the scenario task and gold answer are held fixed across the pair.
  2. Separate measures of participation. Four scored capabilities — fresh-onset response rate, conditional answer accuracy, silence preservation and answering-window yield — are computed on separate denominators, deliberately without forming an aggregate score, because success on one cannot compensate for failure on another.
  3. An automated construction and evaluation pipeline. Generated scripts supply turn labels and gold answers, speech synthesis supplies construction-level audio boundaries, and scoring combines timing checks on the decoded waveform with semantic judgments, with human validation of sampled data and model outputs.

Main Findings

  • Frequent speech does not imply appropriate participation. Under explicit addressing, Freeze-Omni has the highest response presence (0.9955) but the lowest fresh-onset response rate (0.2525), while MiniCPM-o differs by only 0.0045 between presence (0.9545) and fresh onset (0.9500). That gap is exactly the share of requests on which speech was already underway when the request finished.
  • Speech presence says nothing about answer quality. MiniCPM-o and FLM-Audio have similar response presence (0.9545 and 0.9505), yet their conditional answer accuracy is 0.4730 and 0.0011; FLM-Audio produced only 2 correct answers among 1,901 response-present requests.
  • High response rates can coexist with near-total failure to stay silent. MiniCPM-o and Freeze-Omni both have high response presence (0.9545 and 0.9955) but silence preservation of 0.9242 versus 0.0002; Freeze-Omni preserves silence in only 5 of 20,360 silence-requiring windows. Moshi shows the same pattern more mildly: higher response presence than Voila (0.8155 versus 0.6515) but much lower silence preservation (0.2734 versus 0.8351).
  • MiniCPM-o 4.5 leads on three of the four scored capabilities. Under explicit addressing it reports fresh-onset response rate 0.9500, conditional answer accuracy 0.4730 and silence preservation 0.9242. Voila leads answering-window yield under explicit addressing (0.8484) and implicit addressing (0.8046).
  • A transcript-based reference shows a large addressing effect that the speech systems do not. Gemini 3.1 Pro, given speaker-attributed text and the duty instruction, responds at 0.9470 when named and 0.3040 when not — a paired difference of 0.6430 (exact McNemar p < 10⁻³⁷¹). For the five speech systems the differences range from −1.75 to +0.60 percentage points, and all five p values exceed 0.05 (smallest 0.2365), so no systematic response-rate difference between addressing conditions is established.
  • Floor release is a coverage-conditioned outcome. MiniCPM-o produces speech in 1,818 of 1,911 N4 windows, starts speaking within 3 s after the question in 1,736 of them, and is still speaking when N4 R begins in 1,633, which form its answering-window yield denominator. Freeze-Omni has speech in all 1,911 windows, but 1,784 are continuations and only one event reaches the yield denominator. Explicit-condition yield denominators are 1,633 (MiniCPM-o), 219 (Moshi), 977 (FLM-Audio), 620 (Voila) and 1 (Freeze-Omni), with yield estimates of 0.596, 0.192, 0.008 and 0.848 for the first four; rates are reported only when at least 30 events qualify.
  • Stopping failures take two forms. In the explicit condition, FLM-Audio is still speaking 2 s after N4 R begins in 968 of its 977 scored events; in 97 of Moshi's 177 failed events the model is silent at that deadline but speaks again before the observation window ends.
  • Some models respond less as the conversation goes on. Voila's response presence falls from 71.64% to 59.55% (explicit) and 76.21% to 55.82% (implicit) from the earliest to the latest request group; FLM-Audio's fresh-onset rate falls from 67.61% to 47.67% and 66.97% to 48.96%. MiniCPM-o's answer accuracy falls from 50.69% to 46.82% (explicit) and 51.27% to 44.41% (implicit); the other four models stay below 2% in every group.
  • The silence-preservation ranking is robust, but the response metric choice is not. Repeating the analysis with speech-duration thresholds of 0, 100, 200, 300 and 500 ms leaves the ordering unchanged. Ranking the five models' explicit and implicit entries by fresh-onset response rate versus response presence gives a Spearman correlation of −0.4303, and no model keeps the same rank under explicit addressing; Freeze-Omni moves from first to last.
  • Uncertainty checks support the reported ordering. Resampling 2,000 scenarios with replacement and repeating 10,000 times, the 95% intervals support the silence-preservation ordering for all five models under both conditions; for the closest explicit-addressing pair, the interval places MiniCPM-o ahead of Voila by 8.25 to 9.58 percentage points.

Methodology in Plain English

The researchers generated multi-party conversation scripts with Claude Opus 5 from combinations of five scene attributes (setting, activity, relationship, device context and register). Every script contains exactly one still-unresolved request addressed to the assistant, labelled T, plus labelled turns that require silence: N1 (the assistant is mentioned but not asked to act), N2 (the turn is addressed to someone or something else), and N3 (speech with no designated addressee). Some scripts also contain an N4 event, where a participant asks the assistant a question and a later utterance from a human resolves it — the assistant should answer in the roughly 3-second interval between the two and then stop.

Each of the 2,000 scenarios was rewritten into a counterpart that reverses whether the request names the assistant, producing one explicit and one implicit version with the same scenario task and the same gold answer. Qwen3-TTS synthesized every human turn separately under a deterministic speaker-to-voice map, which yielded construction-level gold boundaries.

Five open-weight full-duplex speech systems — MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni — each received a spoken duty preamble followed by continuous room audio, with no transcript, speaker identity, turn boundary or endpoint provided. A scheduler played the human clips in order; a model not already speaking had 5 s after the request to begin a fresh voiced onset, and 3 s of silence were streamed after an N4 question. All timing was measured on the decoded assistant waveform using a causal Silero voice-activity detector. Answer correctness was judged by transcribing responses with Qwen3-ASR-1.7B and having Claude Opus 5 compare meaning against the gold answer; silence preservation allowed a brief acknowledgement that does not take the floor. Six reviewers (two per item) validated both construction and evaluation, with sampled conversations reaching 98%–100% pass rates, all 1,282 checked TTS turns intelligible, all 200 checked ASR transcripts meaning-preserving, and reviewer–Opus agreement of 99/100 on T-answer correctness and 98/100 on acknowledgement-versus-intrusion judgments. Gemini 3.1 Pro was run on speaker-attributed transcripts as a text-only reference — explicitly not an acoustic system or an upper bound.

Why This Matters

Impact on research. The paper shows that a single response rate is an unreliable summary of duplex behaviour: speech presence failed to predict fresh initiation, answer accuracy or silence preservation across all five systems. It also reframes addressee recognition from a label-prediction task into a latent decision that controls what waveform a model emits, and its automated generation-and-scoring pipeline is designed to support larger or fresh evaluation draws rather than a fixed hand-labelled set.

Real-world applications:

  • Meeting assistants that must follow a multi-person discussion and speak only when the contribution is actually requested.
  • Smart speakers and home devices that share a room with other devices and multiple speakers, where wrongly answering a question aimed at a person or another device is a common failure.
  • In-car assistants that must handle overlapping speech from several occupants and stop when a passenger resolves the request themselves.
  • Always-listening wearables and egocentric devices with bystanders present, where deciding not to respond is as important as responding.

Industry relevance. The evaluated systems are all open-weight models with publicly released checkpoints, and the paper publishes precise timing settings, a duty preamble, detector parameters and scoring code, which makes the results directly actionable for teams building or comparing full-duplex products. The finding that no speech system showed a significant explicit-versus-implicit response difference, while a transcript-conditioned reference showed a 64.3-percentage-point one, points to addressee inference from audio as an unresolved engineering gap rather than a solved problem.

Future Directions

  • Closing the addressing gap. Since the speech systems respond at similar rates whether or not the assistant is named, it remains open whether they infer the addressee from context or simply fail to treat the name as an addressing cue; targeted probing could separate these explanations.
  • Improving floor release. Answering-window yield is low for most systems and the eligible-event counts vary enormously, suggesting that training or prompting for stopping behaviour is an open problem, particularly for models that speak almost continuously.
  • Extending beyond the synthetic English scenes. The authors note the scenes are synthetic and do not represent real populations, room acoustics or social norms, and may inherit biases from the generating and speech-synthesis models; more naturalistic audio with real acoustic conditions is a logical next step.
  • Broadening the model set and scenario space. The benchmark currently scores five full-duplex systems, and non-full-duplex models are excluded from scored evaluation because adapting them would require evaluator-chosen invocation points; the authors present instruction sensitivity under a segmented interface only as a non-scored probe.

Target Audience

Researchers and engineers working on full-duplex and streaming speech models, spoken dialogue evaluation, and turn-taking or addressee-recognition systems; developers building multi-user voice assistants for meetings, homes or cars who need to reason about when an assistant should stay quiet; and benchmark designers interested in automated, human-validated construction and scoring pipelines for interactive speech.

Authors’ abstract

Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.

Read the original paper