Skip to content
AI.info

Speech and audio

Voice Conversion and Speech-to-Speech Translation

Distinguish voice conversion, content preservation, style transfer, direct speech translation, cascades, latency, identity, and evaluation.

By the end you can

Example

Eighteen listeners, 151 sentences, and certainty heard in the voice

Certainty is something a listener hears before parsing a single word. In 2018 Xiaoming Jiang and Marc D. Pell put that to a panel. They recorded 151 short English sentences. Each was produced in a scenario in which the speaker was very confident, almost confident but not 100% sure, very unconfident, or neutral. Eighteen native Canadian-English listeners then judged every utterance on a 5-point confidence scale. The agreement was strong enough to use as a filter. A clip was kept as a stimulus only if at least 13 of the 18 listeners — more than 72% — agreed that a level of confidence was conveyed.

Two findings sit on top of each other. Certainty is audible, and it survives an unfamiliar accent. The abstract says so: “We demonstrate that native Canadian-English listeners can recognize confident and doubtful expressions in foreign- and regional-accented speakers. A stronger impression of confidence was shown towards the native speakers.” But who carries the sentence changes how certain it sounds. A linear mixed-effects model found a significant effect of speaker accent on perceived confidence, F(2, 471) = 68.48. The same intended certainty was read differently depending on whose voice delivered it.

That accent, and the prosody wrapped around it, is exactly what a voice conversion or a speech-to-speech translation system replaces. A hedge can leave the source speaker as a hedge and reach the listener as a claim, with every word and every name intact. A review that checks only whether the words and the names came through would pass that output without hesitating. That is the difficulty the rest of the lesson keeps returning to. The evidence anyone thinks to collect is evidence about words. This failure lives entirely in what the words were wrapped in.

  • What has to be decided here runs the full width of the subject: voice conversion, content preservation, style transfer, direct speech translation, cascades, latency, identity, and evaluation.
  • The failure that arrives first is error compounding across a cascade, where each stage hands its mistakes to the next one.
  • The evidence anyone would think to ask for is semantic adequacy and critical-entity preservation — did the meaning survive, and did the names. Neither is sensitive to the accent effect Jiang and Pell measured on perceived confidence, F(2, 471) = 68.48.
  • The practical response is to name the attributes in play before building anything: content, speaker, language, style, timing, emotion, and background.

Case

A hundred languages in, thirty-six out

Before asking what a system preserves, it is worth knowing what it can produce at all. That asymmetry is published. SeamlessM4T, released by Meta in 2023 and reported in Nature in 2024, takes speech input in about a hundred languages. It returns text in about a hundred. It returns speech in thirty-six. The gap is not an oversight. Producing speech needs target-language voice data that text translation does not.

Meta also reports average robustness gains of 37 per cent against background noise and 48 per cent against speaker variation. Both are speech-to-text figures. They sit on the leg of the system that never has to render a voice. Nothing in them speaks to the leg that does — the leg on which prosody is generated, and therefore the leg where the effect Jiang and Pell measured is decided. A pipeline that answers in text for your language and in speech for only a third of them has a product decision inside it. That decision was taken long before anyone wrote down which attributes of a speaker are allowed to survive.

Preserve everything except language is not a brief

Two operations are being asked for here, and they are worth naming apart. Voice conversion changes selected vocal attributes while attempting to preserve linguistic content. Speech-to-speech translation changes language, and may also preserve speaker identity, timing, emotion, or style, through a cascade or through a direct model.

Both attract the same instruction, and the instruction is not a complete brief. “Preserve everything except language” says nothing about the fact that content, identity, emotion, timing, accent, disfluency, uncertainty, background, and turn structure can conflict with one another. Some of them have no direct equivalent in the target language at all.

The Voice Conversion Challenge 2020 measured one of those conflicts across a whole field of systems. It received 33 submissions, 3 of them baselines, and ran crowd-sourced listening tests on all of them. The organizers' result: “In particular, speaker similarity scores of several systems turned out to be as high as target speakers in the intra-lingual semi-parallel VC task. However, we confirmed that none of them have achieved human-level naturalness yet for the same task.” Identity was preserved to the level of the target speaker. Naturalness was not — in the same systems, on the same task. The cross-lingual task, the one that matters for translation, scored lower on both, with the best systems above 4.0 MOS.

So the attributes are not jointly satisfiable. A brief that promises all of them is deciding nothing. Two questions follow it, and they are not the same question. Whether the meaning arrived intact is one. Whether this speaker's voice may carry it at all is the other, and that one needs a signature rather than a score. Neither answer, on its own, should stand in for the complete release decision.

In 33 submissions to the Voice Conversion Challenge 2020, identity reached target-speaker level while naturalness stayed below human on the same task. Leaving the conflicts unnamed hands the tradeoff to whichever attribute the model happened to optimize.

Visual

The contract comes before the architecture

Which is why the order of work is fixed. Write the preservation contract first — the list of attributes that must remain, may change, or must be suppressed. The next decision, cascade or direct model, spends whatever that contract assumed. Evaluating semantics and interaction comes last. It is the point where those assumptions meet an actual conversation.

The four steps are these. Write the preservation contract, stating which attributes must remain, change, or be neutralized. Choose cascade or direct model, comparing ASR–translation–TTS modularity with end-to-end speech translation. Control identity and style through explicit conditioning, consent, and limits on voice preservation. Then evaluate semantics and interaction, testing meaning, names, numbers, prosody, latency, speaker similarity, and repair.

Reverse the order and the evaluation still runs. It simply judges the system against a contract nobody wrote. That is how an output whose words are all correct comes back marked correct while the speaker's perceived confidence has moved — the effect Jiang and Pell sized at F(2, 471) = 68.48, and no word-level check records it.

FigureLayers · 4 layers
  1. 01

    Write the preservation contract

    State which attributes must remain, change, or be neutralized.

  2. 02

    Choose cascade or direct model

    Compare ASR–translation–TTS modularity with end-to-end speech translation.

  3. 03

    Control identity and style

    Use explicit conditioning, consent, and limits on voice preservation.

  4. 04

    Evaluate semantics and interaction

    Test meaning, names, numbers, prosody, latency, speaker similarity, and repair.

Once the preservation contract is written, the semantics and interaction evaluation judges the system against that contract, never against what the contract left out.

Analogy

Restaging a play in another language with the same cast

Writing that contract is a job somebody already does. A play restaged in another language with the same cast keeps each actor's identity and dramatic intent, and loses the gestures and tones that have no equivalent. The director makes those losses deliberately and answers for them afterwards. That is the whole of the difference. A model synthesizes identity and prosody without interpreting anything, and without anyone having granted permission for either.

The permission half of the contract is not only an editorial nicety. In one jurisdiction it is already law, with a date on it. The FCC adopted a Declaratory Ruling on 2 February 2024 and released it on 8 February. FCC 24-17 states: “In this Declaratory Ruling, we confirm that the TCPA’s restrictions on the use of “artificial or prerecorded voice” encompass current AI technologies that generate human voices.” Calls covered by the Telephone Consumer Protection Act therefore require the prior express consent of the called party, absent an emergency purpose or exemption. A later FCC rulemaking, published in the Federal Register on 10 September 2024, refers back to it as "the AI Declaratory Ruling in which the Commission found that AI and other technologies that generate human voices fall within the TCPA".

Read the test carefully. It attaches to generating a human voice at all, not to whether the words were translated correctly. A system with a flawless semantic adequacy score is squarely inside it. Nobody chose to let a hedge sound certain, and nobody signed for the voice either. Both fell out of the architecture and were never recorded as choices at all.

Speech transformation requires an explicit preservation and authorization contract — and since FCC 24-17, released 8 February 2024, the authorization half of it is law in at least one jurisdiction.

Example

Which stage owes the evidence

When that contract becomes a dubbing agreement, it turns on four terms. Used loosely, they leave two things unclear: which stage of the system owes the evidence when something goes wrong, and who may authorize the transformation in the first place.

One of those terms has stopped being an accuracy question. In 2021 a group led by Luisa Bentivogli compared state-of-the-art cascade and direct speech translation systems on English–German, English–Italian and English–Spanish. The comparison used professional post-edits and manual annotation. The paper asks in its title whether the differences between the two paradigms still make a difference, and answers: “Our multi-faceted analysis on one of the few publicly available ST benchmarks attests for the first time that: i) the gap between the two paradigms is now closed, and ii) the subtle differences observed in their behavior are not sufficient for humans neither to distinguish them nor to prefer one over the other.” If annotators can neither tell the two apart nor prefer either, the architecture choice is no longer about output quality. It is about which stages remain open to inspection — that is, about who can answer for the mistake afterwards.

  • Voice conversion transforms vocal identity or style while attempting to preserve linguistic content.
  • A cascade is a system built from separately inspectable stages, which is what makes it possible to ask where the meaning slipped. Since the Bentivogli comparison closed the quality gap, inspectability is most of what it still buys.
  • Direct speech translation generates translated speech without requiring an explicit intermediate transcript in the product interface, which removes the place where that question could have been asked.
  • A preservation contract is the specification of which attributes must remain, which may change, and which must be suppressed.

Example

Where preservation gets negotiated

The contract is not the same document in every setting. Four of them pull it in different directions, and each asks for other units and another kind of proof.

The privacy setting has the sharpest published numbers, and they are humbling. The first VoicePrivacy Challenge drew 45 participants from 13 countries, representing 25 teams, and yielded 16 eligible submissions. Anonymization performance then collapses as the attacker gets better informed. Many systems exceed 50% equal error rate against an ignorant attacker. The best fall to 33–43% against a lazy-informed attacker. Against a semi-informed attacker, whose speaker-verification model has been retrained on anonymized data, they fall to 16–26%. The same converted audio, three different privacy figures, and nothing changed but what the attacker knows. The organizers' own summary, published in 2022: “Anonymization is also achieved only partially and always at the cost of utility; no single system gives the best performance for all metrics and each system offers a different trade-off between privacy and utility, whether judged objectively or subjectively.”

  • In multilingual meetings, translation latency and speaker attribution affect turn-taking.
  • In dubbing, timing and performance need authorized creative review — the director's job, written down.
  • In accessibility, voice conversion can adapt speech while preserving personal identity by choice.
  • In privacy work the aim points the other way, and VoicePrivacy 2020 prices the difference: above 50% equal error rate against an ignorant attacker, 16–26% against a semi-informed one. A stated anonymity level is a property of the assumed attacker rather than of the conversion.

Steps

Write a speech transformation contract

Whichever of the four you are working in, the contract has to be specific enough that another team can use it to demonstrate error compounding across a cascade. It also has to say what inspectability is being bought with, because output quality is no longer what separates the two architectures.

Work through four steps. List the source attributes: content, speaker, language, style, timing, emotion, and background. Choose preservation priorities, resolving conflicts rather than promising all attributes — the Voice Conversion Challenge 2020 is the evidence that they cannot all be had at once. Create contrast cases that test negation, politeness, uncertainty, names, and culture-specific expressions. Then design human review, providing source audio, transcript, translation, and generated output together.

Write down what your preservation contract assumes. Write down one counterexample. A hedge delivered with assertive prosody will do, or an accent swap of the kind that moved perceived confidence by F(2, 471) = 68.48 in the Jiang and Pell panel. Then write down what evaluating semantics and interaction is supposed to trigger when it meets that counterexample.

FigureProcess · 4 steps
  1. 1. List source attributes

    Content, speaker, language, style, timing, emotion, and background.

  2. 2. Choose preservation priorities

    Resolve conflicts rather than promising all attributes.

  3. 3. Create contrast cases

    Test negation, politeness, uncertainty, names, and culture-specific expressions.

  4. 4. Design human review

    Provide source audio, transcript, translation, and generated output together.

Handing over the source audio, transcript, translation and generated output together is what lets a reviewer find the stage where meaning slipped. Any one of them alone hides the compounding.

Example

Did the meaning arrive, and in whose voice

The two questions from earlier become two families of measurement, and neither family carries a release on its own. The first asks whether the meaning arrived. The second asks whether it was allowed to arrive in that voice. Two more complete the portfolio: how errors compound across the stages of a cascade, and whether a user can tell which stage went wrong and repair it.

A community evaluation shows what those families look like once they have units. The IWSLT 2023 evaluation campaign covered 9 shared tasks and attracted 38 submissions from 31 teams: “The shared tasks address 9 scientific challenges in spoken language translation: simultaneous and offline translation, automatic subtitling and dubbing, speech-to-speech translation, multilingual, dialect and low-resource speech translation, and formality control.” Note that dubbing and speech-to-speech translation are scored as separate challenges, not as one. The simultaneous track ran under a single published latency budget: 2 seconds of Average Lagging for speech-to-text, 2.5 seconds of starting offset for speech-to-speech. The offline track drew 10 teams and 37 runs across cascade and direct systems.

That is the standard to copy: a number, a unit, a threshold agreed in advance. Run an output whose words are all correct and whose prosody has moved through the four families. It comes back as a pass on the first line and a failure further down. That combination is the thing the portfolio exists to catch, and no single number can express it.

  • The core task evidence is semantic adequacy and critical-entity preservation.
  • The system behaviour to check is speaker similarity together with explicit consent scope.
  • The robustness slice covers prosody, timing, and interaction latency — which is where the assertive hedge lives, and where IWSLT 2023 shows what a stated budget looks like: 2 seconds of Average Lagging for speech-to-text, 2.5 seconds of starting offset for speech-to-speech.
  • Over the system's working life, the evidence is stage-level error localization and user repair success.

Report semantic adequacy and critical-entity preservation together with stage-level error localization and user repair success, each with a unit and a threshold fixed in advance — the way IWSLT 2023 fixed 2 s Average Lagging and 2.5 s starting offset before anyone submitted.

Key idea

A preserved voice, fluent words, and a $6,000,000 forfeiture

Four things break a confident claim about converted or translated speech: 1) error compounding across a cascade; 2) direct models hiding where meaning changed; 3) prosody transfer misrepresenting certainty or emotion; 4) preserving a speaker's voice without authorization.

The first two are one problem seen from either architecture. A cascade lets errors accumulate but leaves separately inspectable stages to examine. A direct model leaves none. The third is what Jiang and Pell measured. The fourth has a dated, priced instance. The FCC adopted a Forfeiture Order on 26 September 2024 and released it on 30 September. FCC 24-104 imposed a $6,000,000 penalty on political consultant Steve Kramer. What it punished was a January 2024 robocall campaign, two days before the New Hampshire Democratic presidential primary, that used an AI-generated deepfake of President Joseph R. Biden Jr.'s voice with spoofed caller ID. A federal court in New Hampshire put the campaign at approximately 10,000 robocalls and denied the corporate defendants' motion to dismiss. Judge Steven J. McAuliffe, 26 March 2025: “The calls used an AI-generated “deepfake” voice technology to mimic President Biden’s voice and were designed to suppress Democratic voter turnout.”

Nothing in that case turns on a mistranslated word. The voice was preserved, the words were fluent, and the entire harm sat in whose voice was allowed to say them. A usable brief says which of the nine attributes gives way when two of them collide, and which ones the target language simply has no equivalent to hand over. And it says who signed for the voice.

Transferred prosody that turns a hedge into a confident assertion changes what the speaker said, not how they said it, and no translation metric records the difference. A voice used without authorization is not recorded by one either, but FCC 24-104 priced it at $6,000,000.

Key takeaways