Speech and audio
TTS Evaluation, Consent, Provenance, and Deepfake Response
Design TTS evaluation across content, naturalness, intelligibility, similarity, prosody, latency, consent, provenance, misuse, and incident response.
By the end you can
- Define tts evaluation, consent, provenance, and deepfake response as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish content evaluation, perceptual evaluation, and governance evaluation without treating them as interchangeable
- Trace the workflow from define the use and harm model through test governance
- Evaluate tts evaluation, consent, provenance, and deepfake response using text coverage, pronunciation, and critical-entity error and evidence from difficult deployment slices
Ten things, and they move in opposite directions
Text-to-speech quality is not one property that a system has more or less of. It is ten, and each has to be shown separately: content correctness, pronunciation, intelligibility, naturalness, speaker similarity, prosody, latency, accessibility, authorization, and provenance. Nothing holds them together. A system can improve on several of them while getting worse on the rest.
So the work splits three ways. Content evaluation, perceptual evaluation and governance evaluation are run by different people, for different reasons. All three have to be commissioned while the architecture can still change. Two of them ask questions that no amount of later listening will answer.
The rest of this lesson works from four places where that split was tested in public and the record survives. A $6,000,000 forfeiture order over a cloned voice that cost $150 to make. A listening test that gave identical 2013 recordings a different score in 2022. A state statute that put the word “voice” inside a property right. And the EU AI Act, which sets a date after which marking synthetic audio at generation stops being good practice.
Content, perceptual, and governance evaluation belong to different owners, and all three must be commissioned while the architecture can still change.
Example
A $150 clone and a $6,000,000 forfeiture order
A robocall went out to New Hampshire voters in President Biden's voice, telling them not to vote in the primary. The voice was a clone. It cost $150 to make. The forfeiture came to $6,000,000.
The order is FCC 24-104, adopted on 26 September 2024. Its first paragraph says what the recordings were: “Kramer's illegal robocalls carried a deepfake generative artificial intelligence (AI) voice message (Deepfake Message) that imitated U.S. President Joseph R. Biden, Jr.'s voice and encouraged potential voters not to vote in the then-upcoming Primary Election.” The audio was made with Eleven Labs software, the order records, and paid for with $150. Kramer is a political consultant.
Now ask what a listening panel would have contributed. A panel is asked how a voice sounds. Nothing the Commission objected to turns on that — not naturalness, not intelligibility, not prosody, not latency, not long-form stability. The order is about whose voice it was, what the people who answered the phone were told, and whether anybody had agreed to any of it. A five-point rating scale does not contain those three questions and cannot be extended to contain them. A high naturalness score on that audio would have been accurate and beside the point at the same moment.
The same gap opens in the quieter direction. A proper name mispronounced in a narration is a content error, and it does not surface in a naturalness rating either. The listener hears a fluent, pleasant delivery of the wrong word, and rates the delivery. In both directions the mechanism is identical. The instrument answers its own question well and stays silent on the rest.
- This lesson is about one decision: how to design a TTS evaluation that spans all of it — content, naturalness, intelligibility, similarity, prosody, latency, consent, provenance, misuse, and incident response.
- The failure that arrives first is the one in FCC 24-104: audio that works as audio, sitting on top of an authorization question and a disclosure question that no perceptual measurement was ever pointed at.
- The missing evidence sits on both sides of the panel. Text coverage, pronunciation and critical-entity error on one side; consent completeness and provenance coverage on the other. The $150 of synthesis cost was the only figure the process actually produced.
- The practical response is to give content, perception and governance their own acceptance criteria, so that any one of the three can fail without the other two being consulted.
Case
Identical stimuli, a full point lower in 2022
Where does a naturalness rating come from, and what is it a rating of? Mean opinion score comes from ITU-T Recommendation P.800, which defines the five-point absolute category rating scale. P.808 adapts the same procedure to crowdsourced listeners. Both documents specify a listening test in detail: its design, its listeners, its playback, its scale, its anchors, and the question put to them. What comes out is a property of that test. It is not a property the system carries out of the room with it.
That is not a theoretical caution. In 2022 three speech researchers replayed the identical Blizzard Challenge 2013 stimuli to 59 analysed listeners. The scores of the historical systems fell by about a full point: system K from 3.81 to 2.62, system N from 3.22 to 2.00, system C from 2.88 to 1.96. A section headed “How reliable is MOS?” states it without hedging: “Exactly the same stimuli were presented to listeners in both listening tests, yet the MOS scores given by listeners to each and every one of these systems dropped by a full point.” The authors attribute the compression to the modern neural voices sitting in the same test. It might also be a change in the listener population, who are likely more exposed to high-quality synthetic speech than the 2013 listeners. The ranking of the systems was preserved. The same waveforms, a different room, a different number.
So a MOS moves when nothing about the system moves. And neither recommendation says a word about whose voice was used. Somebody else does. On 8 February 2024 the Federal Communications Commission ruled that an AI-generated voice in a robocall is an artificial voice under the Telephone Consumer Protection Act, so the consent rules for artificial and prerecorded voices apply to it. Seven months later the same Commission did the arithmetic in FCC 24-104: 3,000 verified unlawful spoofed calls, a $1,000 base forfeiture each, a 100% upward adjustment, $6,000,000 — against $150 of synthesis. Listening tests and legal exposure are measured on different instruments, and only one of the two has an enforcement arm behind it. A panel score is not a defence. After the Blizzard replay it is not even a stable number.
Comparison
Content, listeners, and authorization
Those instruments belong to a set of three. Content evaluation asks whether the intended words and entities were spoken correctly — text coverage, pronunciation, critical-entity error. A mispronounced name belongs there. Perceptual evaluation asks how the output sounds to listeners under a controlled protocol, and reports naturalness, intelligibility, similarity and prosody. That is the question P.800 and P.808 formalise, and the one the Blizzard replay showed to be a property of the sitting rather than of the system. Governance evaluation asks who was entitled to that voice, for what use, and what the file says about its own origin.
The third column is the one teams treat as internal policy preference. In Tennessee it stopped being one. The ELVIS Act took effect on 1 July 2024, adding “voice” to a property right that until then covered only name, photograph and likeness. The full name is the point of it: the Ensuring Likeness, Voice, and Image Security Act of 2024, in place of the state's Personal Rights Protection Act of 1984. The definition is written so that a synthesised voice cannot slip out of it: “‘Voice’ means a sound in a medium that is readily identifiable and attributable to a particular individual, regardless of whether the sound contains the actual voice or a simulation of the voice of the individual;”. The Act creates civil liability both for publishing an unauthorized voice and for distributing a tool whose primary purpose is producing one. The vendor of the cloning system stands in the same column as its user.
A strong result in one of the three says nothing about the other two. A panel is not run badly when it comes back clean on a voice nobody was entitled to use. It is simply the wrong instrument for two thirds of the question, and it returns a confident number either way.
Content evaluation
Checks whether the intended words and entities were spoken correctly.
- Decision focus: Define the use and harm model
- Useful evidence: Text coverage, pronunciation, and critical-entity error
- Watch for: One listening score hiding content errors
- Best used when its assumptions are documented for tts evaluation, consent, provenance, and deepfake response
Perceptual evaluation
Measures listener judgments under a controlled protocol.
- Decision focus: Build objective checks
- Useful evidence: Naturalness, intelligibility, similarity, and prosody ratings
- Watch for: Listeners recognizing the target speaker and assuming authorization
- Best used when its assumptions are documented for tts evaluation, consent, provenance, and deepfake response
Governance evaluation
Checks authorization, traceability, policy, access, and response mechanisms.
- Decision focus: Run controlled listening
- Useful evidence: Latency, long-form stability, and accessibility outcomes
- Watch for: Provenance added after distribution rather than at generation
- Best used when its assumptions are documented for tts evaluation, consent, provenance, and deepfake response
Key idea
Provenance added after distribution is not provenance
Governance results fail in recognizable ways. Any one of these four conditions puts the release decision outside what listening scores, consent records, and detectors can support.
1) One listening score hiding content errors. 2) Listeners recognizing the target speaker and assuming authorization. 3) Provenance added after distribution rather than at generation. 4) Deepfake response depending on one detector that attackers can adapt to.
The first is the narration case exactly: a fluent delivery of the wrong word rates as a fluent delivery. The second is its mirror image. A listener who knows the voice reads familiarity as permission, which is the one inference a listener is in no position to make. The recipients of Kramer's calls were being invited to make it.
The third has been measured. AudioSeal embeds its watermark at generation and detects it down to a single audio sample — 1/16,000 s. Across audio edits it holds an average ROC AUC of 0.97, where WavMark reaches 0.84. It also localises the manipulated region: “AudioSeal achieves an IoU of 0.99 when just one second of speech is AI-manipulated, compared to WavMark's 0.35.” The after-the-fact alternative does worse on the same material. The passive Voicebox classifier fell to 0.704 accuracy on re-synthesised versus AI-generated audio at 30% masking, with a true positive rate of 0.680 and a false positive rate of 0.194. AudioSeal stayed at 1.0 / 1.0 / 0.0 on that condition. A stamp applied at generation is in every copy that leaves. A stamp applied afterwards reaches only the copies you still have.
The fourth has been measured too, and the numbers are worse still. ASVspoof 5 tuned its spoofing attacks against surrogate countermeasure and speaker-verification systems on purpose. It then ran its two Track 1 baseline detectors, RawNet2 and AASIST, over an evaluation set of 138,688 bona fide and 542,086 spoofed utterances. The set covered 16 attacks unseen in training and development. RawNet2 scored minDCF 0.8266 and an equal error rate of 36.04%. AASIST scored 0.7106 and 29.12%. The paper's own summary of Track 1 is one line: “The baseline systems achieve minDCFs no lower than 0.7 and EERs no lower than 29%.” That is one number standing in for a portfolio, with someone on the other side of it working to make it read the wrong way.
Provenance stamped at generation travels with the file; provenance bolted on after distribution only covers the copies you still control.
Visual
Objective checks before the listening panel
Ten quality dimensions cannot be argued about inside one score, so the order of the work carries as much weight as its content. Four stages stay separable, each checked on its own terms. Define the use and harm model — content, voice, listeners, channel, languages, consequences. Build objective checks — text coverage, pronunciation, critical tokens, duration, technical artifacts. Run controlled listening, with naturalness, intelligibility, similarity, style and preference asked as separate questions, under the protocol P.800 or P.808 specifies. Keep the Blizzard result beside it, a standing reminder that the number belongs to the sitting. Then test governance — consent, provenance, generation logs, abuse controls, revocation, incident response.
The last stage now has a deadline attached to it. Synthetic audio will have to carry a mark a machine can read. Article 50(2) of the EU AI Act: “Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.” Article 50(4) adds a disclosure duty on deployers of deep-fake audio. Both apply from 2 August 2026. Machine-readable marking at generation — the AudioSeal design point — is on a calendar, not on a wish list.
Coming last gives governance testing a limit worth stating plainly. It compares the system against the use and harm model that somebody wrote down. Whatever was never written down has nothing to fail against. The release goes out with a clean sheet.
Define the use and harm model
Specify content, voice, listeners, channel, languages, and consequences.
Build objective checks
Verify text coverage, pronunciation, critical tokens, duration, and technical artifacts.
Run controlled listening
Separate naturalness, intelligibility, similarity, style, and preference questions.
Test governance
Audit consent, provenance, generation logs, abuse controls, revocation, and incident response.
Testing governance only asks whether the system matches the use and harm model you wrote, so harms left out of that model go untested.
Example
MOS is not a governance answer
Four terms circulate through these reviews as if they were grades on the same report card. They are not. They belong to three separate governance conversations, and reading them as one quality score is how a release passes on the strength of the wrong evidence.
- MOS is a mean opinion score obtained from a specified subjective listening protocol — P.800 in the laboratory, P.808 with crowdsourced listeners. The specification is the substance of it. The Blizzard Challenge 2013 stimuli scored 3.81, 3.22 and 2.88 in their own test, and 2.62, 2.00 and 1.96 when replayed to 59 analysed listeners in 2022.
- Speaker similarity is the perceived or modeled resemblance between a generated voice and a reference one. That makes it a fact about acoustics and about nothing else. Tennessee's ELVIS Act reaches the same resemblance as a legal fact, covering a sound attributable to an individual whether it is that person's actual voice or a simulation of it.
- Provenance is traceable information about where an audio artifact came from and what was done to it along the way. From 2 August 2026, under Article 50(2) of the EU AI Act, it is also a machine-readable mark that a provider of a synthetic-audio system is obliged to attach.
- Deepfake response is the set of procedures a team runs when deceptive synthetic media appears: detection, containment, communication, takedown, evidence, and recovery. The detection half returned equal error rates of 36.04% and 29.12% for the ASVspoof 5 Track 1 baselines, against attacks tuned to defeat them.
Example
The critical entity is where to look
A listening score that hides content errors is exposed by one thing only: measuring the content. Text coverage, pronunciation and critical-entity error are where a mispronounced name stops being an impression and becomes a number. In a narration a proper name is a critical entity by definition.
Report every figure with its unit. Report the perceptual figures with the conditions that produced them — who listened, how many of them were analysed, how uncertain the score is, and what they heard it on. The same stimuli moved from 3.81 to 2.62 when the surrounding systems changed. And report the governance figures at the same time as the rest. Those were the only ones with $6,000,000 attached to them.
- For the core task, report text coverage, pronunciation, and critical-entity error.
- For system behaviour, report naturalness, intelligibility, similarity, and prosody ratings, each with the protocol and the listener population that produced them. A MOS without its sitting is the number that fell a full point on unchanged audio.
- For the robustness slice, report latency, long-form stability, and accessibility outcomes, alongside detection performance on attacks the detector has not seen. A17–A32 were unseen in ASVspoof 5 training and development, and the baselines scored minDCF 0.8266 and 0.7106.
- Over the working life of the voice, report consent completeness, provenance coverage, abuse rate, and revocation time. Measure provenance coverage on marks applied at generation — the property that held an average ROC AUC of 0.97 across edits.
Report text coverage, pronunciation, and critical-entity error together with consent completeness, provenance coverage, abuse rate, and revocation time.
Steps
Build a TTS release scorecard
A scorecard has one job: to let another team see, from the record alone, whether a pleasant-sounding voice was carrying content errors or an authorization it never had. Keep it to three things. Write down what your use and harm model assumes. Write down one case that breaks that assumption. Then say what the testing and governance steps do about that case.
Four steps make it operational. First, separate the dimensions and assign independent acceptance criteria to content, perception and governance. Second, sample difficult material — names, numbers, long text, emotion, silence, code-switching. Third, blind the listening test: randomize systems, control level, report listener uncertainty, and record the population, since that population is part of the result. Fourth, run an abuse drill — unauthorized cloning, viral redistribution, takedown, evidence preservation.
The fourth step has a published baseline for how low the bar sits. Consumer Reports assessed six voice-cloning products and published the result on 10 March 2025: Descript, ElevenLabs, Lovo, PlayHT (also branded PlayAI), Resemble AI and Speechify. Only Descript and Resemble AI took technical steps to prevent unauthorized cloning. Of the other four the report says: “Rather, these companies require only that users check a box confirming they have the legal right to clone the voice or make a similar self-attestation.” Even Descript's required spoken consent statement could be defeated by generating that statement with a different voice-cloning service. So the drill is not hypothetical. Ask the same question of the consent artifact in your own pipeline. Is it evidence, or a tickbox that any other vendor's product can produce on demand?
1. Separate dimensions
Assign independent acceptance criteria to content, perception, and governance.
2. Sample difficult material
Use names, numbers, long text, emotion, silence, and code-switching.
3. Blind the listening test
Randomize systems, control level, and report listener uncertainty.
4. Run an abuse drill
Simulate unauthorized cloning, viral redistribution, takedown, and evidence preservation.
Rehearse the misuse chain — unauthorized clone, viral spread, takedown, evidence preservation — before release; a scorecard is worth only the worst case it has already walked through.
Key takeaways
- TTS quality is ten separate things that can move in opposite directions. Any one number reporting it is reporting one of the ten.
- A mean opinion score is a property of a single listening test, not something the system carries out of the room. Replayed to 59 analysed listeners in 2022, the identical Blizzard Challenge 2013 stimuli fell about a full point — system K from 3.81 to 2.62, N from 3.22 to 2.00, C from 2.88 to 1.96. The ranking was preserved.
- Name the intended use and the harms it invites before anything is built. Governance is tested last, and only ever against the model somebody wrote down. Part of that model now has a date on it: Article 50(2) of the EU AI Act requires machine-readable marking of synthetic audio from 2 August 2026.
- Content, perceptual and governance evaluation answer related but different questions, and need acceptance criteria that can fail independently. Where Tennessee's ELVIS Act applies, governance is not an internal preference: its definition of “voice” covers a simulation of an individual's voice as well as the actual one.
- Perceptual quality is no defence against an authorization failure. Forfeiture Order FCC 24-104 imposed $6,000,000 on Steve Kramer over a deepfake robocall imitating President Biden's voice: 3,000 verified spoofed calls at a $1,000 base forfeiture, with a 100% upward adjustment. The synthesis cost $150, on Eleven Labs software.
- Report text coverage, pronunciation and critical-entity error together with consent completeness, provenance coverage, abuse rate and revocation time, and stamp provenance at generation. AudioSeal held an average ROC AUC of 0.97 across edits and an IoU of 0.99 on one manipulated second, while the passive Voicebox classifier fell to 0.704 accuracy at 30% masking.