Speech and audio
Audio Dataset Design, Rights, Consent, and Provenance
Design representative audio datasets with lawful collection, consent, provenance, split integrity, documentation, and lifecycle controls.
By the end you can
- Define audio dataset design, rights, consent, and provenance as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish legal availability, informed consent, and technical provenance without treating them as interchangeable
- Trace the workflow from define purpose and prohibited use through govern the lifecycle
- Evaluate audio dataset design, rights, consent, and provenance using speaker, device, language, room, and time coverage and evidence from difficult deployment slices
Coverage and rights are one design problem
A speech corpus is usually judged the way any sample is judged: how many speakers, how many rooms, how many devices, how many languages, how many hours. But collection source, speaker relationship, license, consent, withdrawal, channel, room, language, device, time and downstream use all sit in the same list. Only some of them are statistics. The rest decide whether the team was allowed to hold the recording at all. No amount of coverage settles them.
The two cannot be separated in audio, because the recording is the person. Public access does not imply unrestricted training rights. It does not imply biometric consent, or permission to generate a person's voice. De-identification, the usual way out, barely works either. Strip the name from the metadata and the voice itself is still there. So is the conversation happening behind it, where the speaker was, how they sounded when unwell, and who they were talking to. Identity, background conversation, location, health cues, social context — all of it is still in the waveform. A bigger model does not help here.
So the question that decides a corpus is not how broad it is. It is whether the purpose and the prohibited uses declared at collection survive into how the dataset is later maintained, corrected and retired. Without that link, a collection can be broad and balanced and still be data the team was never permitted to use. The rest of this lesson shows what that costs, through a federal judge, a federal regulator, a vendor's own apology and a controlled experiment. The gap gets priced.
A corpus that fails on rights is unusable no matter how well its speakers, rooms, and devices are covered.
Case
Read audiobooks, donated sentences, two consent stories
Two of the most-used English corpora answer that question in opposite ways, and the useful thing is that both are public. LibriSpeech takes roughly 1,000 hours of read English from public-domain audiobooks and ships under CC BY 4.0. Its rights come from the source material and the license attached to it, settled before the corpus existed. Common Voice inherits permission from nothing. It asks people to donate recordings and to validate other people's. Consent there is not a document that arrived with the audio. It is the collection mechanism itself. Its 2020 corpus paper reports 2,500 hours across 29 languages from more than 50,000 contributors.
LibriSpeech is worth reading for a second reason. Its authors wrote down what split integrity costs to build. The 2015 paper states the requirement plainly: “In order to guarantee that there was no speaker overlap between the training, development and test sets, we wanted to ensure that each recording is unambiguously attributable to a single speaker.” Meeting it meant excluding multi-reader LibriVox genres. It meant checking chapters with the LIUM diarization toolkit. It also meant capping how much of any one voice a set may contain. dev-clean and test-clean each hold 40 speakers, 20 male and 20 female, at about 8 minutes per speaker — roughly 5.4 hours. dev-other and test-other hold 33 speakers at about 10 minutes each. The training subsets are capped too: 25 minutes per speaker in train-clean-100 (100.6 h, 251 speakers) and train-clean-360 (363.6 h, 921 speakers), 30 minutes in train-other-500 (496.7 h, 1,166 speakers). Somebody counted the minutes, voice by voice, before anyone trained on them.
Two defensible corpora, defensible for different reasons. Neither is a general-purpose corpus. Read audiobooks and donated sentences are not the same speech, and neither one is a call centre. What they share is that somebody wrote down where the recordings came from and on what terms — and, for LibriSpeech, who exactly is on each one. Early enough for the answer to still be true.
Figure
Example
Recordings bought on Fiverr for $1,200 and $400, then trained into a product
A third team built a corpus the way neither of those was built, and a federal judge has now described how. Two voice actors sold recordings on Fiverr, for $1,200 and $400. The stated purposes were "academic research purposes only" and "test scripts for radio ads". Lovo employees bought them and used them to train Genny, the company's commercial text-to-speech model. On 10 July 2025, in Lehrman v. Lovo, Judge J. Paul Oetken ruled that Paul Lehrman and Linnea Sage could proceed on New York Civil Rights Law §§ 50–51, consumer-protection and breach-of-contract claims. Most of the federal copyright and Lanham Act claims were dismissed. What survived is the part about what the speakers were told.
What they were told is on the record. A Fiverr message reached Linnea Sage from the user "tomlsg", alleged to be Lovo co-founder Tom Lee. It reads: “These are test scripts for radio ads. They will not be disclosed externally, and will only be consumed internally, so will not require rights of any sort.” The audio was obtained lawfully. It was paid for. It was delivered with a purpose attached. The purpose is exactly what the training use walked past.
Report only how broadly the recordings cover speakers, devices, rooms, and times, and none of that appears in the numbers at all.
- The decision this lesson is about is the one Lovo never made: design an audio dataset that is representative and lawfully collected, settling consent, provenance, split integrity, documentation and lifecycle controls together rather than one at a time.
- The failure that quietly does the most damage is narrower than the consent problem, and much easier to miss — the same speaker appearing on both sides of the train and test split. The next sections put a number on it.
- The evidence anyone will ask to see is coverage: speaker, device, language, room and time. That is exactly the evidence which stays silent about a Fiverr message promising the recording would not be disclosed externally.
- The practical response is to link every recording to its source, its rights, its consent, the transformations applied to it and the annotators who have touched it. Then any one of those can be answered later without guessing — including, at $1,200 and $400 a clip, what purpose the seller was told.
Comparison
Licensed, consented, traceable: pick all three
Put the corpora side by side and the questions separate cleanly. Legal availability asks whether the recording can be accessed or licensed under stated terms. CC BY 4.0 answers it for LibriSpeech. A paid Fiverr order answered it for the recordings Lovo bought. Informed consent asks something narrower: whether the person recorded agreed to this particular use. Common Voice puts that question at the front of its collection. Linnea Sage was answered on it in advance, wrongly: “These are test scripts for radio ads. They will not be disclosed externally, and will only be consumed internally, so will not require rights of any sort.” Technical provenance asks what neither of the first two touches. Can you still show, later, where a given recording came from and what has been done to it since? LibriSpeech answers that with diarization checks and per-speaker minute caps. A purchased clip folded into a training set answers it with nothing.
What answers one leaves the other open. A permissive license does not produce consent. Payment does not produce consent either: the §§ 50–51, consumer-protection and breach-of-contract claims went forward against a company that had paid. Consent given once does not produce a record of what happened afterwards. Legal availability means the recording can be accessed or licensed under stated terms. It means nothing beyond that.
Legal availability
The recording can be accessed or licensed under stated terms.
- Decision focus: Define purpose and prohibited use
- Useful evidence: Speaker, device, language, room, and time coverage
- Watch for: Speaker leakage across train and test splits
- Best used when its assumptions are documented for audio dataset design, rights, consent, and provenance
Informed consent
A participant understands the specific collection and downstream use.
- Decision focus: Map rights and participants
- Useful evidence: Consent and license completeness by release
- Watch for: Bystander speech collected without meaningful notice
- Best used when its assumptions are documented for audio dataset design, rights, consent, and provenance
Technical provenance
The system can trace source, edits, annotations, transformations, and release history.
- Decision focus: Design representative sampling
- Useful evidence: Duplicate and near-duplicate leakage rate
- Watch for: Licenses that exclude model training or synthetic derivatives
- Best used when its assumptions are documented for audio dataset design, rights, consent, and provenance
Visual
The step that gets dropped is the people
Provenance is the question a process is supposed to answer, and it sits at a specific place in the sequence. A dataset process opens by defining the purpose and the prohibited uses. It closes by governing the lifecycle — corrections, removals, retirement. Between the two sits the step that maps rights and participants: who is on these recordings, on what terms, having agreed to what.
That is the step that gets dropped. It gets dropped because it is the only one that requires finding something out rather than writing something down. Drop it, and whatever the purpose statement assumed about the people on the tape arrives at lifecycle governance unchecked. Lovo was not missing a document. The purpose was stated in writing, on Fiverr, to the person being recorded. What was missing was the step where someone checks the training use against that statement before the model ships. So the check happened in a courtroom instead, on 10 July 2025.
1. Define purpose and prohibited use
State the intended model, decision, users, and uses the dataset must not support.
2. Map rights and participants
Track speakers, bystanders, creators, licensors, locations, and jurisdictions.
3. Design representative sampling
Cover acoustic, linguistic, demographic, device, and temporal variation relevant to deployment.
4. Govern the lifecycle
Version releases, honor withdrawal, restrict access, document lineage, and define retention and deletion.
Whatever purpose and prohibited uses were written down travel into lifecycle governance as settled fact, and nothing later reopens them.
Key idea
The same speaker on both sides of the split
Nothing in the four failures below looks like misconduct at the moment it happens. That is why they do the most damage — to a dataset, and to the people recorded in it.
1) Speaker leakage across train and test splits. 2) Bystander speech collected without meaningful notice. 3) Licenses that exclude model training or synthetic derivatives. 4) Dataset documentation becoming stale after corrections and removals.
The first is a measurement failure, and it is the reason a corpus can look strong and be hollow. Split a body of recordings without tracking who is speaking and the test set stops being a test. That is the failure LibriSpeech spent genre exclusions and diarization checks to prevent, and the one the evaluation section prices. The third is Lovo precisely: a purpose assumed rather than confirmed. The audio was paid for, so training on it, enrolling the voice and regenerating it got treated as a single right. The fourth is slower. The corrections and removals happen, the documentation stops describing them, and the description outlives the dataset it describes.
The second is the hardest, because the bystander is not in the metadata to begin with. It also has a dated, self-documented instance. Contractors grading Siri regularly heard confidential material; The Guardian reported it on 26 July 2019. Apple suspended human grading worldwide and apologised on 28 August 2019. Its own statement describes the programme: “Before we suspended grading, our process involved reviewing a small sample of audio from Siri requests — less than 0.2 percent — and their computer-generated transcripts, to measure how well Siri was responding and to improve its reliability.” The same statement says “As a result of our review, we realize we haven't been fully living up to our high ideals, and for that we apologize.” It commits that “Our team will work to delete any recording which is determined to be an inadvertent trigger of Siri.” By default Apple would no longer retain audio recordings of Siri interactions, review became opt-in only, and only Apple employees would listen.
Note what that last commitment implies. The recordings nobody meant to make had to be found after the fact and deleted. Less than 0.2 percent was enough to force it. And deleting a name from the metadata rescues none of it, for the reason this lesson opened with: what remains is a voice, a room, a health cue and a conversation.
Speaker leakage turns a test set into a second training set, and the number it produces will hold up right until the model meets someone it has never heard.
Steps
Create a dataset release dossier
Documentation going stale is the fourth of those failures, so the test of a dossier is not whether it exists. It is whether anyone else can use it. It is only a dossier when another team can pick it up and check, without you, whether a voice in the test set was already in the training set. That is the first failure, caught by a stranger. A release note written in your own words will never do that.
You do not have to invent the format. Datasheets for Datasets, published in 2018 by Timnit Gebru and six colleagues, asks for one. Communications of the ACM ran it in 2021. The analogy is an electronic component's datasheet, and the proposal is one line of the abstract: “By analogy, we propose that every dataset be accompanied with a datasheet that documents its motivation, composition, collection process, recommended uses, and so on.” The published sections cover motivation, composition, collection process, recommended uses, distribution and maintenance. Collection process and maintenance are the two an audio corpus cannot afford to leave blank.
Three things go on the page beyond that template. Write down what the declared purpose and prohibited uses assume about who is on the recordings and what they agreed to. Add one counterexample: a use that fits the stated purpose and is still wrong, the way training Genny fitted a paid Fiverr order for radio-ad test scripts. Then say what the lifecycle rules require when that counterexample turns up.
1. Build a provenance table
Link each recording to source, rights, consent, transformations, and annotator history.
2. Test split integrity
Detect repeated speakers, sessions, devices, scripts, and near-duplicate clips.
3. Run a misuse review
Consider impersonation, surveillance, health inference, and unintended population use.
4. Set retirement triggers
Define what forces restriction, correction, reconsent, or deletion.
Until the dossier names what triggers restriction, correction, reconsent, or deletion, it documents one day rather than rules anyone can hold you to.
Example
Withdrawal latency is a number you owe
A counterexample is cheap to write down and expensive to honour. The evaluation is where that shows. On 31 May 2023 the FTC and DOJ sued Amazon over children's Alexa voice recordings. The complaint alleged that Amazon retained them indefinitely, and that transcripts were not deleted from all its databases even when parents asked. Amazon agreed to a $25 million civil penalty, to delete inactive children's Alexa accounts, and to a prohibition on using deletion-requested voice data "for the creation or improvement of any data product". Samuel Levine, Director of the FTC's Bureau of Consumer Protection, put the principle in one line: “COPPA does not allow companies to keep children’s data forever for any reason, and certainly not to train their algorithms.” Amazon's own statement reads "While we disagree with the FTC's claims and deny violating the law, this settlement puts the matter behind us." That is what deletion verification costs when someone else measures it. The request was made. The deletion was not shown. $25 million and a ban on using the data followed.
The other conflict is arithmetic. Coverage will sometimes fight withdrawal latency, deletion verification and access incidents, for the obvious reason: honouring a withdrawal removes a speaker, and removing a speaker costs coverage. That is where the evaluation changes the decision instead of decorating it. Which is why the second set of numbers belongs beside the first, not in a separate document nobody reads.
And the abstain threshold now has something to be set against. A 2026 study measured what speaker leakage is worth. On the DAIC-WOZ depression corpus, researchers held training-set size constant and varied only speaker overlap. The split runs over 189 subjects and 6,545 speech segments: a control group of 151 subjects with 5,117 segments, a target group of 38 subjects with 1,428 segments. Their introduction reports the result: “Across architectures, performance improves dramatically under speaker-overlapped evaluation but drops sharply under strict speaker independence (e.g., 97.65% to 58.74% accuracy for a fine-tuned Wav2Vec model).” Same corpus, same amount of training data. 97.65% against 58.74% is the whole of the gap between a model that recognises depression and one that recognises voices. Agree the threshold before the measurement arrives, not after.
- The core task evidence is coverage: how many speakers, devices, languages, rooms and times the corpus actually spans — the figure LibriSpeech publishes down to 40 speakers and about 8 minutes each in test-clean.
- For system behaviour, report consent and license completeness release by release, so a shortfall shows up as a trend across releases rather than as a single awkward line.
- The robustness slice is the duplicate and near-duplicate leakage rate, which is speaker leakage measured rather than assumed away. On DAIC-WOZ, unmeasured, it was worth 97.65% minus 58.74% of headline accuracy.
- The lifecycle evidence is withdrawal latency, deletion verification and access incidents: how long a withdrawal takes to take effect, whether a deletion was confirmed rather than merely requested — the exact gap priced at $25 million on 31 May 2023 — and who reached the recordings who should not have.
Report speaker, device, language, room, and time coverage together with withdrawal latency, deletion verification, and access incidents.
Key takeaways
- LibriSpeech's roughly 1,000 hours of public-domain audiobook reading under CC BY 4.0 and Common Voice's 2,500 hours across 29 languages from more than 50,000 contributors are both public and both defensible — one by license, one by asking. The recordings Lovo bought on Fiverr for $1,200 and $400 came with a stated purpose that was neither.
- Payment is not permission. On 10 July 2025, in Lehrman v. Lovo, Judge J. Paul Oetken let §§ 50–51, consumer-protection and breach-of-contract claims proceed against a company that had paid for the audio, because the seller had been told “These are test scripts for radio ads. They will not be disclosed externally, and will only be consumed internally, so will not require rights of any sort.”
- A dataset is defensible only when the purpose and prohibited uses declared at collection still govern how it is later maintained, corrected and retired. The step that breaks that link is the one mapping rights and participants, and it is dropped because it is the only step that requires finding something out.
- Legal availability, informed consent and technical provenance answer related but different questions: a license does not produce consent, and consent does not produce a record of what happened next. That is why Datasheets for Datasets asks every dataset to ship with a datasheet covering motivation, composition, collection process, recommended uses, distribution and maintenance.
- Splits that place the same speaker on both sides inflate scores. On DAIC-WOZ, with training-set size held constant, a fine-tuned Wav2Vec model scored 97.65% under speaker-overlapped evaluation and 58.74% under strict speaker independence. Decide before you measure what level forces you to abstain.
- Report speaker, device, language, room and time coverage together with withdrawal latency, deletion verification and access incidents. Amazon paid a $25 million civil penalty and accepted a ban on using deletion-requested voice data "for the creation or improvement of any data product" because deletions were requested and not shown to have happened.