Speech and audio
Voice Biometrics, Privacy, and De-Identification
Cover voice templates, linkability, membership and attribute inference, de-identification, pseudonymization, utility tradeoffs, and revocation.
By the end you can
- Define voice biometrics, privacy, and de-identification as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish pseudonymization, voice transformation, and feature release without treating them as interchangeable
- Trace the workflow from define the protected unit through evaluate privacy and utility
- Evaluate voice biometrics, privacy, and de-identification using closed- and open-set re-identification under strong attackers and evidence from difficult deployment slices
Comparison
Each defends against a different attacker
Three ways of protecting a voice are usually on the table. They are not three versions of one idea. Pseudonymization replaces the direct identifiers and keeps linkability alive on purpose, through a controlled key: two recordings can still be recognised as the same speaker by whoever holds it. Voice transformation changes the acoustic identity cues and keeps selected content. Feature release shares embeddings or derived features that may still permit linkage or inference. Each gives up something different in exchange for what it protects.
Which of the three is right is not a question about the audio. Each has to be defended against a different attacker. Until someone says who that attacker is and what they already know, none of the three can be shown to work at all. The field has a standing answer to that question. The VoicePrivacy Challenge scores anonymisation against adversaries of declared knowledge — ignorant, lazy-informed and semi-informed — rather than against a hope. Run the same systems past all three and the numbers move. They move a long way.
Pseudonymization
Replaces direct identifiers while preserving linkability through a controlled key.
- Decision focus: Define the protected unit
- Useful evidence: Closed- and open-set re-identification under strong attackers
- Watch for: Evaluating privacy against only one weak attacker
- Best used when its assumptions are documented for voice biometrics, privacy, and de-identification
Voice transformation
Changes acoustic identity cues while retaining selected content.
- Decision focus: Specify the attacker
- Useful evidence: Linkability across sessions and releases
- Watch for: Consistent pseudovoices enabling cross-session tracking
- Best used when its assumptions are documented for voice biometrics, privacy, and de-identification
Feature release
Shares embeddings or derived features that may still permit linkage or inference.
- Decision focus: Transform and minimize
- Useful evidence: Intelligibility, content, and downstream utility preservation
- Watch for: Lexical content revealing identity despite acoustic transformation
- Best used when its assumptions are documented for voice biometrics, privacy, and de-identification
Example
The same systems: above 50%, then 33–43%, then 16–26%
Skipping that question costs something, and the VoicePrivacy 2020 Challenge printed the bill. The submitted anonymisation systems were scored more than once, against attack models of stated knowledge. Between those scorings the audio did not change at all. Only the attacker did.
The spread is the whole lesson. Against the ignorant attacker, many systems reached equal error rates above 50%. Against the lazy-informed attacker, the best results were 33–43%. Against the semi-informed attacker, the same systems came in at only 16–26%. That attacker's automatic speaker verification model had been retrained on anonymised speech. A team reporting the first number and a team reporting the third are describing identical audio.
The organisers stated the consequence plainly when they published the results in 2022: “Thus, assessing the performance of anonymization systems using an ASV system trained on original data leads to a false impression of protection.” False impression is the operative phrase. Nothing had gone wrong with the transforms. The protection reported at the top of the range was manufactured by the test, not by the audio.
- What is being decided here spans the whole subject: voice templates, linkability, membership and attribute inference, de-identification, pseudonymization, the utility a transform trades away, and revocation.
- The failure that arrives first, almost everywhere, is evaluating privacy against only one weak attacker. It is the distance between an EER above 50% and one of 16–26%, on unchanged speech.
- The evidence that would have caught it is closed- and open-set re-identification measured under strong attackers — in the 2020 results, an ASV system retrained on anonymised speech rather than on original data.
- The practical response is to write down what the protected unit actually includes: the waveform, the transcript, the embedding, the metadata, the timestamps and the model outputs.
Case
The organisers retired their own weak attackers; the statutes arrived separately
The organisers' answer to that spread was not to argue about it. They retired the weak rungs. The 2022 evaluation plan kept two things: unprotected speech, and a strengthened semi-informed attacker. That attacker is an automatic speaker verification system retrained on utterance-level anonymised LibriSpeech-train-clean-360. Of it the plan says: “This attack model is actually the strongest known to date, hence we consider it as the most reliable for privacy assessment.” Equal error rate under that attacker is the privacy metric. Word error rate is the primary utility metric. The protection and its cost are reported in the same breath.
Read back onto the 2020 results, the arrangement is sharp. A system scored at the bottom of that ladder and then described as anonymisation is making a claim about the top of it. The two numbers involved are above 50% and 16–26%, on identical audio.
Meanwhile the law does not wait for the metric to settle. Illinois lists a voiceprint as a biometric identifier in its Biometric Information Privacy Act, and that statute requires a written release before a private entity collects one. Texas enforces its own version at scale. On 9 May 2025 the Texas Attorney General announced a $1.375 billion settlement in principle with Google. It resolves 2022 suits that alleged, among other things, capture of biometric identifiers including voiceprints through Google Assistant, Google Photos and Nest Hub Max, in violation of the Texas Capture or Use of Biometric Identifier Act. It is the largest state privacy recovery against Google to date. “For years, Google secretly tracked people’s movements, private searches, and even their voiceprints and facial geometry through their products and services,” said Attorney General Ken Paxton. A voiceprint statute is not a paper obligation waiting for an EER to be agreed.
A transformed voice is not an anonymous one
Why does the transform not settle it by itself? Because identity in speech is not kept in one place. Speech is linkable through vocal, linguistic, behavioral and contextual cues. Voice de-identification changes selected identity evidence while attempting to preserve content or utility. Selected is the load-bearing word. What a method leaves alone is exactly what a retrained adversary works with.
So the privacy that results is never a property the transform can hand over. The 2020 systems were fixed objects. Their measured protection ran from above 50% down to 16–26% because the knowledge on the other side changed. Privacy depends on what the attacker knows, what auxiliary data exists, how often the same speakers are released, and what the downstream task is. That is why the challenge attaches a named adversary to every number it reports. It is also why its organisers bothered to rank their own attackers and call the last one the strongest known to date.
There is a second boundary, and it runs the other way. Perfect unlinkability may conflict with intelligibility, speaker consistency, emotion, pathology or forensic utility — the properties the audio was kept for in the first place. That is the trade word error rate sits in the score to measure, beside equal error rate. Defensible practice therefore ties one end to the other: define the protected unit at the start, evaluate privacy and utility at the end, with a named owner at each end.
Anonymity is a claim about what an attacker already knows, not a property the transform can hand over: the same systems scored above 50% and 16–26% on audio that never changed.
Example
The words a privacy argument is lost in
Naming the unit and naming the adversary both depend on four words that are routinely swapped for one another. Swap one for another and you quietly change what evidence is owed, what unit it is measured in, and who holds the decision. The four have to be kept apart on purpose.
- Linkability is the ability to determine that two records belong to the same person or source. It is what the semi-informed attacker restored well enough to drive the 2020 systems down to 16–26% EER.
- Pseudonymization replaces identifiers while retaining a controlled means of linkage. Of the three approaches, it is the one that keeps linkage available deliberately.
- A voice template is a stored representation used for speaker comparison, and it can outlive the recording it was derived from. That is why HM Revenue & Customs had to count enrolments, not recordings, when the Information Commissioner's Office came for them.
- Attribute inference is estimating sensitive characteristics from released audio or representations. It is a leak that lands even when nobody is ever re-identified.
Visual
Name the attacker before the transform
Those four words belong to different stages of the same job, and the job goes wrong when the stages arrive as one number. Privacy work fails quietly when measurement, modeling, decision and verification are collapsed into a single score. The score reads well. Nobody can say which stage produced it — or, as in the 2020 results, which attacker produced it.
The path below pulls those stages apart and takes them one at a time, in the order the challenge implies. Say who the attacker is and what unit is being protected before choosing the transform, not after reporting on it. The order matters because the 2022 plan shows what a late correction costs. The attacker was rebuilt, and every number measured against the old one stopped meaning what it had meant.
1. Define the protected unit
Choose recording, utterance, session, speaker, household, or cohort.
2. Specify the attacker
Document enrollment data, auxiliary recordings, models, and cross-release access.
3. Transform and minimize
Use only necessary audio, remove metadata, and apply voice conversion or representation controls.
4. Evaluate privacy and utility
Measure re-identification, linkability, attribute leakage, intelligibility, and task performance.
Privacy and utility get measured against the protected unit you named earlier, so naming the wrong unit still produces numbers that look reassuring.
Key idea
Amazon deleted the recordings and kept the transcripts
Taken in the wrong order, four failures follow. Each of them raises the reported number rather than lowering it, which is why none of them looks like a failure from inside the results.
The first is evaluating privacy against only one weak attacker — the 2020 spread, and the reason the organisers climbed their own ladder and then threw the lower rungs away. The second is consistent pseudovoices enabling cross-session tracking: a pseudonym stable enough to be useful within a session is stable enough to be followed between sessions. The third is lexical content revealing identity despite acoustic transformation, the habit of word choice that no pitch shift touches. The fourth is deleting raw audio while retaining irreversible biometric templates and derived artefacts indefinitely. That one retires the evidence and keeps the exposure.
The fourth is not hypothetical. On 31 May 2023 the FTC and DOJ filed a complaint and proposed order against Amazon. It alleged that Amazon retained children's Alexa voice recordings indefinitely, and that the deletion parents asked for did not reach everything derived from them. The agency's own summary: “And even when a parent sought to delete that information, the FTC said, Amazon failed to delete transcripts of what kids said from all its databases.” Amazon agreed to a $25 million civil penalty, to delete the retained data, and to stop using it to train its algorithms. The waveform was one item in the protected unit. The transcript was another. The second one is the reason the deletion did not hold.
The second failure in the list is hard rather than careless. The trade described earlier cuts both ways: in some applications a voice that holds steady is the point. Which is why the trade has to be settled before release rather than after it.
Deleting the recording is not deleting the record — the FTC's 2023 complaint against Amazon turned on transcripts that survived the parents' deletion requests, and cost $25 million.
Example
Four places a voice stays identifying
Where that trade gets settled depends on what the audio was kept for. Four settings differ enough that no single ruling covers them. Research corpora, call analytics, voice authentication and synthetic voices each leak identity in their own way, and each is counted in its own unit. Nothing that closed the 2020 evaluation closes any of them.
- In a research corpus, transformed speech still requires access and purpose controls. The transform lowers the risk; it does not license open release. The challenge systems were measured, not cleared.
- In call analytics, transcripts and speaking patterns can remain identifying long after the voice has stopped sounding like anyone. The transcripts Amazon kept were the record that mattered, not the audio it deleted.
- In voice authentication, template compromise is harder to remedy than a password leak, because the credential cannot be reissued. And the enrolments outlive the sessions that created them, as HM Revenue & Customs found when it had to account for every one.
- In synthetic voices, a consistent pseudonym is what makes interaction work and what makes long-term tracking possible. It is one property doing both jobs.
Example
Other teams cut the published EER by 25-44%
Across all four settings the report has the same shape. Closed- and open-set re-identification is only informative when the attacker is the strongest one the team can build. And the strongest one available is a moving figure, not a fixed setting.
The First VoicePrivacy Attacker Challenge made that concrete in 2025. Participants built attacker automatic speaker verification systems against anonymisation systems already submitted to VoicePrivacy 2024. In the authors' words, “The best attacker systems reduced the equal error rate (EER) by 25-44% relative w.r.t. the baseline.” The speech had not changed. The anonymisation systems had not changed. Published privacy figures fell by a quarter to nearly a half because other people were invited to attack them.
So report the unit the measure is in. Then say whose voices were tested, how uncertain the result is, and which attacker it was tested against, beside linkability across sessions and releases. A privacy figure with no adversary named beside it is not yet a privacy figure. One measured against last year's adversary has an expiry date.
- For the core task, closed- and open-set re-identification under strong attackers — retrained on anonymised speech, as the 2022 plan requires, not on original data.
- For system behavior, linkability across sessions and releases — the consistent pseudovoice, measured.
- For the robustness slice, whether intelligibility, content and downstream utility survived, which is the word-error-rate half of the same trade.
- Over the working life of the release, attribute leakage, revocation success, and access-control incidents — plus a re-test when someone stronger shows up, because 25-44% of the reported protection went that way once already.
A privacy number is a claim about one adversary on one date: in 2025, outside teams cut the baseline attacker's equal error rate by 25-44% without touching the speech.
Steps
Write a voice privacy threat model
A threat model is only worth writing if it names the adversaries the authors would rather not think about. That is the difference between a published ladder and a team's own reassurance. Because the ladder states what each attacker knows, another team can use it to challenge a privacy claim that was only ever tested at the bottom of it.
Start by writing down what defining the protected unit assumes: which of the waveform, transcript, embedding, metadata, timestamps and model outputs is in scope, and how far the release travels. Then give one counterexample to those assumptions — an adversary who has retrained on your anonymised output, or a second release of the same speakers. Then say what evaluating privacy and utility does with that counterexample: which number it moves, in which unit, and against which named adversary it was measured.
The last step, deletion and revocation, is the one that gets written as a sentence and tested as an operation. After a Big Brother Watch complaint, the UK Information Commissioner's Office moved to force HM Revenue & Customs to delete Voice ID biometric enrolments collected without explicit consent. A preliminary enforcement notice came on 4 April 2019, a final one on 9 May 2019. HMRC's chief executive, Sir Jonathan Thompson, wrote in a public letter dated 3 May 2019: “I have confirmed that HMRC will only retain Voice ID enrolments where we hold explicit consent.” Behind that sentence is arithmetic. Around 1.5 million customers had enrolled since the October 2018 changes, and those enrolments were kept. The roughly 5 million records without explicit consent — enrolments taken before October 2018 — were already being deleted, and the work was to finish well before the ICO's 5 June 2019 deadline. Five million voiceprints, one deadline, one letter naming who signed for it. That is what a finished revocation clause looks like when someone makes you execute it.
1. Map released artifacts
Include waveform, transcript, embedding, metadata, timestamps, and model outputs.
2. List auxiliary data
Assume public media, leaked enrollment, and repeated observations.
3. Choose privacy goals
Separate unlinkability, attribute hiding, content secrecy, and access control.
4. Set deletion and revocation
Define how users remove recordings, templates, derived artifacts, and future use.
No threat model is finished until it states how a person gets their recordings, templates, and derived artifacts removed — HMRC deleted roughly 5 million Voice ID enrolments before a 5 June 2019 deadline and kept around 1.5 million.
Key takeaways
- In the VoicePrivacy 2020 Challenge, the same anonymisation systems reached equal error rates above 50% against the ignorant attacker, 33–43% at best against the lazy-informed attacker, and only 16–26% against the semi-informed attacker retrained on anonymised speech. The audio was identical in all three scorings.
- The organisers' own conclusion is that “assessing the performance of anonymization systems using an ASV system trained on original data leads to a false impression of protection”. So the 2022 evaluation plan kept only unprotected speech and a strengthened semi-informed attacker they call “the strongest known to date”. EER is the privacy metric, WER the primary utility metric.
- The strongest attacker you can build is a moving figure. At the First VoicePrivacy Attacker Challenge in 2025, outside teams reduced the equal error rate by 25-44% relative to the organisers' baseline attacker, on anonymisation systems that had not changed.
- A transformed voice is not automatically anonymous. Speech stays linkable through vocal, linguistic, behavioral and contextual cues, and what a method leaves untouched is what a retrained adversary works with. Pseudonymization, voice transformation and feature release each have to be defended against a different attacker.
- De-identification starts by naming what is being protected — waveform, transcript, embedding, metadata, timestamps, model outputs — because deletion reaches only what was named. The FTC and DOJ complaint against Amazon of 31 May 2023 alleged that transcripts of what children said survived their parents' deletion requests. Amazon agreed to a $25 million civil penalty.
- Revocation is an operation, not a clause. After a Big Brother Watch complaint, HM Revenue & Customs deleted roughly 5 million Voice ID enrolments held without explicit consent ahead of the ICO's 5 June 2019 deadline, keeping around 1.5 million. And voiceprint statutes are enforced: Illinois lists a voiceprint in its Biometric Information Privacy Act and requires a written release, and Texas announced a $1.375 billion settlement in principle with Google on 9 May 2025.