Skip to content
AI.info

Speech and audio

Multi-Speaker, Multilingual, and Zero-Shot Voice Cloning

Cover speaker embeddings, reference encoders, few-shot adaptation, zero-shot cloning, language transfer, disentanglement, misuse, and consent.

By the end you can

Example

9,581 calls in a voice the speaker never authorized

On 21 January 2024, 9,581 calls reached New Hampshire telephones. Every one of them carried an AI-cloned voice of President Biden. All 9,581 had been initiated by 7:12 p.m. that evening. The FCC's forfeiture order in the matter records that the message was made with ElevenLabs. The voice was recognisable. That was the point of it, and it was also the whole of the wrong: the person the voice belonged to had authorized nothing.

A review that asked only how closely the cloned voice matched the original would have passed this without hesitating. The match was excellent. Matching was the point. That is the difficulty the rest of the lesson keeps returning to. The measurement everyone thinks to take is the one measurement that cannot see this failure. And the amount of audio needed to produce it is now very small indeed.

  • What has to be decided here runs the full width of the subject: speaker embeddings, reference encoders, few-shot adaptation, zero-shot cloning, language transfer, disentanglement, misuse, and consent.
  • The failure that arrives first is the one on display above — cloning a person's voice without their specific consent. Here it ended in a $6,000,000 penalty, adopted on 26 September 2024.
  • The evidence anyone would think to ask for is speaker similarity, established with independent human and model evidence. On this incident it would have read as a success.
  • The practical response is to stop recordings gathered for transcription, authentication or support from feeding a synthesis system automatically.

Case

Three seconds of audio, and a statute that says a simulation counts

Three seconds is not a figure of speech. Microsoft's VALL-E, described in 2023, synthesises an unseen speaker's voice from a three-second enrolled recording. It does not update the model to do it. That is what makes each additional voice essentially free. The cost was paid up front instead, in scale: pre-training on 60,000 hours of English speech from LibriLight, hundreds of times more than earlier systems used. Nowhere in that pipeline is there a step at which consent is recorded, checked, or withdrawn.

Legislatures noticed. Tennessee's ELVIS Act — the Ensuring Likeness Voice and Image Security Act — was signed on 21 March 2024 and took effect on 1 July 2024. The enacted text is more specific than the summaries of it. It writes a definition into state law: “"Voice" means a sound in a medium that is readily identifiable and attributable to a particular individual, regardless of whether the sound contains the actual voice or a simulation of the voice of the individual”. It then gives every individual a property right in the use of that individual's name, photograph, voice, or likeness, in any medium and in any manner. And it reaches past the clip to the tool. Distributing an algorithm, software or other technology whose primary purpose is the production of an individual's voice without authorization is itself a violation.

Put the robocall between those two facts and the gap is easy to see. The statute does not ask how close the match was. It asks whose voice it is attributable to. It counts a simulation as that voice, and holds the answer as property. A similarity score is evidence about acoustics. What the statute protects is not acoustics.

Position

Consent arrives with the capability, not after it

The usual order is the reverse of that. Zero-shot cloning is demonstrated as a capability, and the permission question is handed on to whoever handles permissions. The dates make that order hard to defend. The capability was published in 2023. The regulators arrived in 2024, by which time the enrolled recordings and the systems built on them already existed.

In the weeks after the New Hampshire calls, the Federal Communications Commission held that a voice cloned or generated with AI is an "artificial" voice under the Telephone Consumer Protection Act. It adopted that ruling on 2 February 2024 and released it on 8 February. The operative sentence is a consent requirement: “Therefore, callers must obtain prior express consent from the called party before making a call that utilizes artificial or prerecorded voice simulated or generated through AI technology.” Note which consent that is — the called party's. The impersonated speaker's authorization is a second permission, and the ruling does not supply it. The ELVIS Act is where that one lives. A cloned voice can be wrong twice over. Neither wrong is acoustic.

No acoustic measurement can supply permission. That is why similarity numbers are never the whole record. Consent coverage, provenance, misuse reports and revocation completion belong beside them, and the last of the four decides whether the other three mean anything. Reference audio, embeddings, adapters and caches all persist. A permission that cannot be withdrawn from every one of them was never a permission. It was a licence taken once, from somebody who had been recorded for something else.

Build the revocation path before the demo, not after the first complaint.

Analogy

A highly skilled impersonator with a digital studio

The nearest human equivalent shows what is and is not new. An impersonator can reproduce a voice after hearing a short sample, and so can a model. In neither case does the resemblance grant permission to speak for that person. The difference is what happens next. The impersonator works one performance at a time, and can be held to account for that performance. A model generates unlimited content, distributes it at scale, and combines identities with nobody deciding to. No impersonator has ever delivered 9,581 performances in a single evening.

An impersonator also stops when asked. That is the part with no counterpart in a system whose reference audio, embeddings and adapters have already been copied somewhere else.

Voice similarity is a capability; authorization is a separate governance decision.

Similarity is not ownership

Underneath both the impersonator and the model sits the same mechanism. Multi-speaker text-to-speech conditions generation on a speaker identity — an embedding, a reference recording, or an adapter. Zero-shot cloning is the version of that which synthesises a new speaker from limited reference audio without updating all model parameters. Which is exactly why three seconds and no training run were enough.

The operational boundary matters as much as the mechanism, and it is a short sentence: speaker similarity does not establish consent, authenticity, or exclusive ownership. Those are records a person holds, not quantities computed from audio. In Tennessee they are a statutory property right, since 1 July 2024. On a phone call they are prior express consent under the Telephone Consumer Protection Act, since February 2024.

So a build that can be defended starts by deciding which identity uses are allowed, and finishes by evaluating and governing those same uses. With a named owner at each end.

Skip the named owner at either end and the cloned voice ships with nobody answerable for how it gets used.

Example

A consent form that blurs these permits more

Four different things happen inside that build, and a consent form that runs them together permits more than it means to. Each names different evidence. Each implies a different person entitled to authorize the use.

  • Speaker conditioning is the information supplied to control which speaker identity gets synthesized.
  • A reference encoder is the model that summarizes a reference utterance so synthesis can be conditioned on it — the component that turns a few seconds of somebody's recorded speech into something reusable.
  • Zero-shot cloning is generating an unseen speaker from a short reference without speaker-specific full retraining. That is the three-second VALL-E case, and the reason an extra voice costs almost nothing to add.
  • Disentanglement is the attempt to represent factors such as speaker, content, and prosody separately.

Key idea

The room the reference was recorded in became part of the voice

A consent form is not the only thing that can be looser than intended. So can the reference recording — and the cloning paper reports that as a feature.

Start with the numbers. On the LibriSpeech test-clean evaluation, VALL-E scores a WavLM-TDNN speaker similarity of 0.580. The YourTTS baseline scores 0.337, and ground-truth speech 0.754. The word error rates run 5.9, 7.7 and 2.2. That is what "excellent similarity" actually looks like as a number: a large gain over the baseline, still well short of the real speaker. The same paper reports that the model preserves the speaker's emotion and the acoustic environment of the three-second prompt. The room travels with the voice by design. The similarity score books that as success.

Four problems recur, and a similarity score looks excellent through every one of them.

1) Training or cloning without specific consent. 2) Reference noise or room becoming part of the cloned identity. 3) Cross-language output caricaturing accent or pronunciation. 4) Revocation failing because embeddings and derivative models persist.

The middle two are where the score is most thoroughly blind. Accent, health cues, emotion, background artifacts and demographic stereotypes can all be reproduced or exaggerated while the identity still registers as a match. The first and the fourth are the same problem at opposite ends of the system's life, and neither of them is a measurement at all. The paper's own broader-impacts paragraph names the first: “Since VALL-E could synthesize speech that maintains speaker identity, it may carry potential risks in misuse of the model, such as spoofing voice identification or impersonating a specific speaker.” The warning shipped with the capability. The consent step did not.

Ask for consent at the moment of cloning, not after: once embeddings and derivative models have spread, withdrawing it is an engineering problem no similarity score will flag.

Visual

Is the output identity the authorized one

That fourth problem is a sequencing problem, and the sequence is short enough to hold in the head. Deciding which identity uses are allowed comes first. Building the speaker conditioning comes second. Evaluating and governing those uses comes last. The middle step is the one that gets skipped, and skipping it leaves a single assumption untested: that the identity in the output is the identity that was authorized. For the 9,581 calls of 21 January 2024, that assumption was false from the first synthesized word. No stage of the pipeline was positioned to notice.

FigureTimeline · 4 stops
  1. 1. Define allowed identity use

    Specify enrollment, consent, languages, styles, duration, revocation, and prohibited content.

  2. 2. Build speaker conditioning

    Use IDs, embeddings, reference encoders, adapters, or prompt audio.

  3. 3. Separate factors cautiously

    Model content, speaker, language, prosody, and channel without assuming perfect disentanglement.

  4. 4. Evaluate and govern

    Test similarity, intelligibility, cross-language behavior, misuse, provenance, and deletion.

Everything the governance review checks rests on the identity uses you allowed earlier, and the review never goes back to that list.

Example

One capability, and one of its uses is fraud

One capability points in four directions at once. One of those directions has a file number.

The New Hampshire calls ended in a $6,000,000 penalty. The FCC's forfeiture order, adopted 26 September 2024, opens: “We impose a penalty of $6,000,000 against Steve Kramer (Kramer) for effectuating an illegal robocall campaign that targeted potential New Hampshire voters two days before the state’s 2024 Democratic Presidential Primary Election (Primary Election) in violation of the Truth in Caller ID Act of 2009”. The arithmetic behind that figure is worth reading, because it shows what enforcement can actually reach. A $1,000 base forfeiture per call, doubled to $2,000, multiplied across 3,000 verified calls — a third of the 9,581 that went out. And a penalty is not a control. On 13 June 2025 a Belknap County Superior Court jury acquitted Kramer of all 22 state criminal counts.

The biometric half of the risk has a public measurement. ASVspoof 5, published in 2024, built its evaluation set from the English partition of Multilingual LibriSpeech: more than 4,000 speakers recorded in non-studio conditions. Its attacks were generated with pre-trained YourTTS and XTTS — the same class of openly available systems this lesson has been describing. Established countermeasures did not survive the move off studio audio: “The baseline systems achieve minDCFs no lower than 0.7 and EERs no lower than 29%.” Only the best five of 54 participants' submissions stayed below 15%.

  • In assistive communication, users may preserve or create a personalized synthetic voice.
  • In localization, authorized performers can produce multilingual variants with review — and cross-language output is precisely where accent and pronunciation get caricatured.
  • In creative tools, character voices need rights, credits, and output controls behind them.
  • Fraud runs on the identical capability. The same open systems that serve the three uses above appear in ASVspoof 5 as the attacks, and baseline countermeasures sat at equal error rates no lower than 29%.

Example

Excellent similarity, failing revocation

Fraud is the reason the report cannot be a single number. Speaker similarity, backed by independent human and model evidence, can be excellent at the very moment consent coverage, provenance, misuse reports and revocation completion are all failing. Catching that combination is what the evaluation exists for. The New Hampshire robocall would have appeared in such a report as a success on the first line and a failure on the last.

It is worth knowing what the human half of "independent human and model evidence" is actually worth. A 2023 study in PLOS ONE played genuine and deepfake audio to 529 people, in English and in Mandarin: “Listeners only correctly spotted the deepfakes 73% of the time, and there was no difference in detectability between the two languages.” Training the listeners with examples improved that only slightly. Human evidence is a real instrument with a known error rate, not a ground truth.

The reference-duration slice has published numbers too. The YourTTS paper, in 2022, reports: “Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality.” Its speaker-adaptation experiment used Common Voice samples of 20 to 61 seconds from four speakers. It concluded that more than 45 seconds is needed for higher quality. So the honest version of the duration bullet is two thresholds, not one. Similarity arrives in under a minute. Quality wants more than 45 seconds. Neither threshold is long enough for anybody to notice they were being enrolled.

The same portfolio has to mark the cases where consent is missing or too narrow. Those are the ones where the team abstains or falls back rather than synthesizes.

  • For the core task, speaker similarity established with independent human and model evidence — reported with the human half's own error rate, which was 73% correct across 529 listeners in two languages.
  • For system behaviour, intelligibility, naturalness, and pronunciation reported language by language; VALL-E pairs its 0.580 similarity with a word error rate of 5.9 against 2.2 for ground truth, and both belong in the report.
  • For the robustness slice, sensitivity to reference duration and to channel — under a minute for similarity, more than 45 seconds before quality holds, and recorded through what.
  • Over the working life of the voice, consent coverage, provenance, misuse reports, and revocation completion.

Report speaker similarity with independent human and model evidence together with consent coverage, provenance, misuse reports, and revocation completion.

Steps

Design a consent-aware cloning workflow

A cloning workflow is consent-aware only if another team can inspect it and point to the step that would have caught the robocall before anything was dialled. Four steps carry that weight. Separate purposes, so transcription, authentication and support recordings never feed synthesis automatically. Verify enrollment against a real identity and an explicit authorization for defined output uses. Control generation with content policy, rate limits, watermark or provenance, and audit logs. And support revocation across reference audio, embeddings, adapters, caches and future access.

The fourth step is the one that has already been litigated. On 31 May 2023 the FTC and DOJ filed a complaint against Amazon. The agencies alleged that it had retained children's Alexa voice recordings indefinitely. They alleged it had failed to delete transcripts from all databases when parents requested deletion, and had used the retained recordings to improve its speech algorithms. Amazon agreed to a $25 million civil penalty and a court order. “COPPA does not allow companies to keep children’s data forever for any reason, and certainly not to train their algorithms,” said Samuel Levine, who directs the FTC's Bureau of Consumer Protection. The order requires deletion of inactive child accounts, and bars using data subject to deletion requests to create or improve any data product. That is the anatomy of failure mode 4 at company scale: a delete button that reached the recording, missed the transcripts, and left the model that had already learned from both.

Write down what your definition of allowed identity use assumes. Then write one case that breaks that assumption, and say what the evaluation and governance steps do about it — including which database the deletion request does not currently reach.

FigureProcess · 4 steps
  1. 1. Separate purposes

    Do not reuse transcription, authentication, or support recordings for synthesis automatically.

  2. 2. Verify enrollment

    Confirm identity and explicit authorization for defined output uses.

  3. 3. Control generation

    Add content policy, rate limits, watermark or provenance, and audit logs.

  4. 4. Support revocation

    Delete reference audio, embeddings, adapters, caches, and future access.

Deletion is the test of the whole design — if a withdrawal cannot reach reference audio, embeddings, adapters, caches and future access alike, the consent step was decorative.

Key takeaways