Skip to content
AI.info

The Pulse

Qwen Cuts Live Translation Lag to 2.3 Seconds

Qwen has released Qwen3.8-LiveTranslate, a real-time interpretation model that reduces average lagging from 2.8 seconds to 2.3 seconds. The model adds speaker separation, synchronized bilingual output, long-context disambiguation, and suppo

Qwen Cuts Live Translation Lag to 2.3 Seconds

AI.info Team ·

Qwen released Qwen3.8-LiveTranslate on September 18, presenting a real-time interpretation model that cuts average lagging from 2.8 seconds in the previous generation to 2.3 seconds. The company says the model improves translation faithfulness, fluency, and conciseness while adding speaker separation and synchronized source-and-translation output.

The release targets live conversations rather than ordinary turn-based translation. Qwen says the system can identify who is speaking, preserve each speaker’s vocal character in translated speech, and display the original and translated language together. The model supports audio input with text output in 60 languages and translated audio output in 29.

Qwen describes the release as Qwen3.8-LiveTranslate, with a real-time service model listed as qwen3.8-livetranslate-flash-realtime.

Qwen Reworks Translation Around an Interleaved Stream

The model uses what Qwen calls an Interleave architecture, which combines audio and text in a single streaming sequence. The company says previously heard audio and already generated translation can be cached and reused, allowing the system to maintain context while reducing delay.

Its architecture uses two modules: a Hybrid-MoE-based Thinker and Talker. The Thinker processes video, audio, source text, and translation in temporal order. The Talker then combines the translation with the source audio to generate speech intended to preserve the original speaker’s timbre.

That design differs from a simple chain of speech recognition, text translation, and speech synthesis. Qwen presents the system as an end-to-end interpretation pipeline in which understanding, translation, and speech generation remain connected during the live exchange.

Speaker Separation Becomes a Product Feature

Qwen3.8-LiveTranslate adds real-time speaker separation for conversations involving multiple people. The system attributes sentences to different speakers and uses that information when producing translated speech, which Qwen says makes voice cloning more stable and helps listeners follow who said what.

The model also synchronizes the source language and translation on one bilingual display. Qwen frames the feature as useful for immediate comprehension and source checking, while also pointing to applications such as subtitles, content organization, and retrieval.

Long-context disambiguation handles another common problem in live interpretation: names, references, and terms whose meaning depends on earlier turns. Qwen says the model links current speech with prior text and conversation history to keep those expressions more consistent.

Claims Span 70 Evaluation Directions

Qwen evaluated the system on the Omnilingua-MSpeaker multi-speaker long-audio set, which covers 14 language directions. The company says Qwen3.8-LiveTranslate outperformed mainstream real-time interpretation systems on translation faithfulness, fluency, conciseness, and diarization error rate.

For broader multilingual testing, Qwen used the public FLEURS audio test set across 70 language directions. The company says the new model led both Qwen’s previous generation and current mainstream systems on translation quality, average lagging, speech-recognition accuracy, and speech-synthesis quality.

The announcement does not publish the underlying benchmark tables in its text, so the release’s comparisons are best read as claims from Qwen rather than independently verified rankings.

DashScope Provides the First Deployment Path

Qwen’s example for using the model runs through the DashScope API over a WebSocket connection. A client streams microphone audio to the service and receives translated text, synthesized audio, or both, depending on the session configuration.

The same interface can return speaker identifiers and source-language text alongside the translation. Qwen’s example also supports sending image frames as visual context, allowing the service to use information beyond the audio stream when interpreting an exchange.

The release lists 60 supported languages for audio input and text output, including English, Chinese, Arabic, French, German, Hindi, Japanese, Korean, Spanish, Turkish, Ukrainian, and Vietnamese. Audio output covers 29 languages, including English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Arabic, Hindi, and Persian.

Qwen says its next development targets are lower end-to-end delay, memory that carries across sessions, and expanded coverage for less widely supported languages and regional dialects. For the current release, the measurable change is narrower: average lagging falls by half a second, to 2.3 seconds, while the service adds speaker attribution and bilingual synchronized output.

Source

Qwen

Explore

More articles