The Pulse
DeepL Voice Preserves Speakers’ Voices During Live Translation
DeepL has introduced voice models that preserve each speaker’s voice, tone and delivery during real-time translation across 12 languages. The company also launched a Windows and Mac desktop app for Zoom, Microsoft Teams and Google Meet, wit

AI.info Team ·
DeepL Voice Adds Speaker Identity to Live Translation
DeepL says its latest Voice models can preserve each speaker’s distinct voice and delivery while translating speech between languages in real time. The company announced the update on September 15, 2026, alongside a new desktop application that brings the same Voice experience to Zoom, Microsoft Teams and Google Meet.
The change targets a weakness in live speech translation: a system can convey the words while flattening the person who said them. DeepL says its models retain differences in tone, rhythm, pacing and intonation, allowing participants to recognize who is speaking even when several voices move through different languages. The models also aim to carry across questions, excitement, hesitation, emphasis and urgency.
DeepL describes the release in its announcement as an expansion across its Voice products, including online meetings, in-person conversations and the DeepL API for Voice.
Twelve Languages, With More Planned
Voice preservation is initially available in 12 languages. DeepL does not list the full set in the announcement, but says additional languages will follow.
The feature is designed to operate without storing voice samples or creating voice profiles for later use, according to the company. That distinction matters for business deployments in which meeting audio, speaker identity and personal voice data may fall under internal security rules or privacy obligations. DeepL presents the processing model as real-time voice transformation rather than a system that retains a reusable profile of each participant.
DeepL’s announcement places the new voice capability on top of an existing translation system. The company cites independent testing by Slator that measured a 4% translation error rate for DeepL Voice, compared with an average of 17% for Microsoft Teams, Google Meet and Zoom. The figures are presented by DeepL as part of the case for extending the product from word accuracy into delivery and expression.
One Desktop App for Zoom, Teams and Meet
The new DeepL Voice desktop app runs on Windows and Mac and works alongside three major meeting services: Zoom, Microsoft Teams and Google Meet. DeepL says the app can detect meetings automatically and start translation with one click, rather than requiring a separate integration for each conferencing platform.
Participants can follow a conversation through translated subtitles in a customizable overlay, voice translation, or both. The app operates outside the individual meeting platforms, which allows DeepL to add features without waiting for Zoom, Microsoft or Google to update their own products.
Voice-to-voice translation for online meetings through the desktop app is now generally available. That capability joins voice translation for in-person conversations and the DeepL API for Voice, which lets organizations integrate speech translation into their own software and services.
Browser Translation Remains in Beta
DeepL is drawing a line between the new desktop release and its browser-based experience. Voice-to-voice translation for external participants in a browser remains in beta, according to the company’s announcement.
The distinction gives the desktop application a more defined role for organizations that control the meeting setup or distribute software to employees. Browser-based access may be easier for outside participants, but DeepL is not presenting that route as finished. The company’s general-availability claim applies to voice-to-voice translation through the new desktop app, not every version of the feature.
Translation That Keeps the Speaker Recognizable
DeepL’s product decision reflects a shift in what users expect from live translation. Captions can make a meeting understandable, but a translated voice that preserves a speaker’s manner may help participants track the conversation more naturally, especially when several people are talking or when the meaning depends on urgency and emphasis.
The release also broadens DeepL Voice beyond a single meeting workflow. Its coverage now spans online meetings, face-to-face conversations and an API, while the desktop application gives the company one distribution layer across three conferencing services. For users, the concrete change is simple: translated speech is intended to sound less like a generic system voice and more like the person who originally spoke.