Speech and audio
Speech and Audio AI as Measurement and Decision Engineering
Frame speech and audio AI as a chain from physical sound through measurement, representation, inference, and product action.
By the end you can
- Define speech and audio ai as measurement and decision engineering as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish signal task, interaction task, and decision system without treating them as interchangeable
- Trace the workflow from define the acoustic event through bind output to action
- Evaluate speech and audio ai as measurement and decision engineering using task accuracy on representative acoustic slices and evidence from difficult deployment slices
Example
Seven point four errors per hundred words, and a signed note
A speech recogniser in a hospital got 7.4 words wrong in every hundred. The notes it produced were reviewed, signed by a physician, and acted on by other clinicians. That is the distance this lesson is about: between a transcript that looks correct and an action a professional puts their name to.
The count comes from 217 clinical notes dictated by 144 physicians in 2016. Zhou and colleagues annotated every error in them and published the result in JAMA Network Open in 2018. Three sentences of the abstract carry the whole chain: "The error rate in SR notes was 7.4% (ie, 7.4 errors per 100 words). It decreased to 0.4% after transcriptionist review and 0.3% in SNs. Overall, 96.3% of SR notes, 58.1% of MT notes, and 42.4% of SNs contained errors."
Read that as a system rather than as a model score. The recogniser's raw output was wrong 7.4 times per 100 words. Human transcriptionists absorbed most of it and took the rate to 0.4. The signing physicians took it to 0.3. And 42.4% of the signed notes still contained errors, of which 6.4% were clinically significant — against 5.7% at the raw stage and 8.9% after transcriptionist editing.
A review that scored only the recogniser would have reported one number. It would have missed the two stages of human correction that number silently assumed. It would have missed the residual that reached a patient record anyway. The dictation chain will keep coming back in this lesson. It is the shortest example of every point that follows.
- What is at stake is seeing speech and audio AI as one chain — physical sound, then measurement, then representation, then inference, then human correction, and only at the end a product action — rather than as a model with plumbing bolted to it.
- The failure that starts most of the others is confusing a waveform with an objective record of the world.
- The evidence that would have caught a 7.4-per-100-word raw rate is task accuracy measured on audio that looks like the audio the system actually meets, reported stage by stage rather than pooled.
- The first practical move is to describe the physical or communicative event in plain words — a physician dictating a note that another clinician will act on — without naming a model at all.
No model has ever heard a room
The recogniser in that hospital did the narrow thing about as well as its conditions allowed. It is worth being precise about what the narrow thing was, because it was not hearing. An audio model never receives "sound itself." It receives measurements produced by microphones, codecs, clocks, channel layouts, preprocessing, and product choices. Between the physician's voice and the words in the note sat a stack of equipment decisions. Not one of them was made with the clinical consequence of an error in mind.
So reliable design starts somewhere other than model selection. Four things get written down first: the acoustic event, the observation process that turns that event into numbers, the output, and the decision that follows from the output. Which model you choose matters less than whether the event you defined is still the event the decision turns on.
Where that link holds, an accuracy figure measured on audio like the deployed audio can support the decision. Where it breaks, the number is still true and still irrelevant. It breaks the moment a 7.4% raw rate is quoted without saying that two stages of human review stand between it and the signed note.
Every recording carries the fingerprints of the equipment that made it, so a score earned on someone else's gear is not evidence about yours.
Case
Eighty-six systems, twelve datasets, two numbers
The measurement half of that chain is now taken seriously enough that the field's own leaderboards have stopped reporting a single figure. The Open ASR Leaderboard states its scope exactly: "It compares 86 open-source and proprietary systems across 12 datasets, with English short- and long-form and multilingual short-form tracks." It prints word error rate beside inverse real-time factor, so accuracy and speed have to be read together. It declines to pool its audio into one score.
The headline count is itself a moving number. The paper first appeared in October 2025. The current version is dated 30 March 2026. Earlier versions reported 60+ systems across 11 datasets. The 86 and the 12 belong to a version stamp, not to the field. That is this lesson's argument arriving inside its own citation: a figure travels with its conditions, and one of those conditions is the date on the document you read it in.
The scale underneath the rankings is genuine, and it is a claim about supervision rather than about audio quality. Whisper's own paper put it this way in 2022: "When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning."
Two numbers over twelve datasets is a great deal more honest than one. 680,000 hours is a great deal of audio. Neither tells you whether a transcript should have been signed into a patient record.
Key idea
Treating confidence as permission to act
Between a leaderboard number and an executed action sit four shortcuts. Each was defensible to whoever took it. Each leaves the product to absorb something the model was never asked to handle.
The first is taking a waveform for an objective record of the world. Five commercial systems — Amazon, Apple, Google, IBM, Microsoft — were run over 19.8 hours of interviews with 42 white and 73 black speakers across five US cities. The finding, published in PNAS in 2020: "We found that all five ASR systems exhibited substantial racial disparities, with an average word error rate (WER) of 0.35 for black speakers compared with 0.19 for white speakers." The gap persisted on a subset of identical phrases spoken by both groups. That is the part that matters here. It was the slice, not the sentence, carrying the result.
The second shortcut treats model confidence as authorization to act, as though a high score were a permission slip. The third evaluates isolated clips and then deploys on a continuous stream. The fourth ignores the people who are recorded — or synthesised — but never became users. Steve Kramer's robocall campaign is the fourth in its purest form. The recipients were the ones the audio acted on, and none of them were anybody's user.
Playback at the microphone is no longer an unmeasured hypothetical either. ASVspoof 5 built a whole benchmark for it in 2024. Its speech came from MLS English — "more than 4k speakers, recorded with diverse devices" — against roughly 100 anechoic-chamber speakers in earlier editions. It carries 32 crowdsourced attack algorithms, A01 to A32, and the 16 attacks A17–A32 form the evaluation set. Adversarial attacks appear for the first time. Across 54 challenge participants the organisers' summary is blunt: "Attacks significantly compromise the baseline systems, while submissions bring substantial improvements."
What the four shortcuts have in common is a number carried past its conditions. Curated clips certify curated clips. Then the audio runs continuously, the room is unfamiliar, the device is new, the speaker is not the speaker the benchmark sampled, a recording is played back at the microphone, or the two directions of error cost different amounts. At that point the benchmark is being asked to cover a case it never measured.
A recording documents one room, one device, and one moment; read as a neutral record of the world, it pushes risk out of the model and into the product.
Example
The same chain under four different loads
The chain that produced a 7.4-per-100-word raw dictation rate — microphone, model, output, correction, decision — is the same chain everywhere else. What changes from one product to the next is where the load falls on it, and with it the evidence that would justify a release. Four settings show the range. No two of them are settled by the same measurement, or reported in the same units.
- Meeting assistance looks like one task and is four: transcription, diarization, summaries and action extraction each demand their own evidence.
- Industrial monitoring has to catch faults it has never heard. DCASE 2020 Challenge Task 2 set the first public benchmark for unsupervised anomalous sound detection, on subsets of the ToyADMOS and MIMII datasets, and stated the difficulty honestly: "The main challenge of this task is to detect unknown anomalous sounds under the condition that only normal sound samples have been provided as training data." It drew 117 submissions from 40 teams. On the evaluation set the official baseline reached 82.80% average AUC (65.80% pAUC), against 94.54% AUC (84.30% pAUC) for the top-ranked submission, Giri_Amazon_task2_2. Which is why the defensible action is to send someone to inspect the machine, not to stop it automatically.
- A voice interface cannot be judged on recognition quality alone. That quality has to be combined with turn-taking and a confirmation policy — precisely the piece that two stages of transcriptionist and physician review were supplying by hand in the dictation chain.
- Media workflows span classification, search, editing and provenance, and those four carry different units and different risks.
Example
Say waveform, not recording
The argument so far has rested on four words, and they only carry it if they are used exactly. Reach for a near-synonym — "recording" for waveform, "clip" for inference unit — and three things change quietly along with it: the proof owed, the unit it comes in, and the person entitled to decide. It is the same discipline that makes 0.35 against 0.19 a usable number and "substantial racial disparities" only a summary of it. A document is challengeable only while its terms stay fixed.
- A waveform is a sampled sequence representing changes in pressure, or in some other sensor quantity, over time.
- An acoustic scene is everything that produced the recording: the physical sources, the room, the noise, the propagation, and the arrangement of the sensors.
- An inference unit is whatever a single prediction actually applies to — a frame, a clip, an utterance, a turn, an event or an entire session. 7.4 errors per 100 words and 42.4% of notes containing errors are two different units over the same audio.
- Abstention is a deliberate refusal to make or execute a prediction when the evidence is not sufficient.
Analogy
The audio evidence chain
A courtroom is the same chain in formal dress. Evidence is transformed four times on its way to a ruling — by a microphone, a stenographer, an analyst and a judge. A clean transcript at the second stage can still support a bad ruling when context or authority is missing at the fourth. That is the dictation chain restated in older language.
One difference is worth holding on to. Testimony is given once and thereafter remembered. Recorded audio can be copied, remixed and replayed at a scale no courtroom was ever built to handle. Synthesised audio removes even the original utterance from the chain.
On 30 September 2024 the FCC released Forfeiture Order FCC 24-104: "We impose a penalty of $6,000,000 against Steve Kramer (Kramer) for effectuating an illegal robocall campaign that targeted potential New Hampshire voters two days before the state's 2024 Democratic Presidential Primary Election (Primary Election)". The campaign used an AI deepfake of President Joseph R. Biden, Jr.'s voice, with spoofed caller ID, to tell voters not to vote.
Then note what two forums did with the same audio artefact. On 13 June 2025 a Belknap County jury acquitted Kramer of all state criminal charges — 11 felony voter-suppression counts and 11 candidate-impersonation counts. A $6,000,000 federal forfeiture and a full state acquittal, over one recording. The artefact settles nothing about the decision built on it.
Accuracy at one stage cannot substitute for a defensible end-to-end decision.
Steps
Write an audio decision contract
This becomes one document, the audio decision contract, and the document has a single test. Hand it to another team and it should give them enough to argue that your recording was never the neutral record of the room you took it for. If they cannot mount that argument from what you wrote, you have not written enough.
Four things go in. What your definition of the acoustic event takes for granted — name the event before you name a model. What the measurement path does to it: sensor, room, channel, preprocessing, storage. One concrete case where the assumption fails: a speaker whose group the benchmark under-sampled at 0.35 against 0.19, a second voice, a recording played back at the microphone in the style of A17–A32. And what the rule linking output to action does when it does fail: abstain, fall back, or put the decision in front of a person, as the transcriptionist and the signing physician did for 7.4 errors per 100 words.
One clause is not yours to design. On 8 February 2024 the FCC ruled that the Telephone Consumer Protection Act's restrictions on "artificial or prerecorded voice" reach AI technologies that generate human voices. Declaratory Ruling FCC 24-17 puts the consequence in one sentence: "Therefore, callers must obtain prior express consent from the called party before making a call that utilizes artificial or prerecorded voice simulated or generated through AI technology." Absent an emergency purpose or exemption, that is a precondition on the action. No accuracy figure moves it.
1. Name the event
Describe the physical or communicative event without naming a model.
2. Draw the measurement path
List sensor, room, channel, preprocessing, and storage transformations.
3. Define the action
State what changes after a prediction and who can stop it.
4. List invalid evidence
Record what a transcript, score, or embedding cannot prove.
Until the contract names what a transcript, score, or embedding cannot prove, it describes the model rather than binding anyone to a rule.
Example
Accuracy is one of four numbers
The Open ASR Leaderboard reports two numbers over 12 datasets. A release needs four, and the contract is what tells you which four. Task accuracy is only the first. On its own it is exactly the figure that would have cleared a dictation pipeline whose raw output was wrong 7.4 times per 100 words.
Held together, the four stop accuracy from carrying a release claim it was never able to carry alone.
- The core evidence is task accuracy measured on acoustic slices that represent the audio the system will actually meet, reported per slice — 0.35 for one group of speakers and 0.19 for another is two numbers, and their average is neither.
- System behaviour is latency and real-time factor at the product boundary, where somebody is waiting for an answer. Inverse real-time factor sits beside word error rate on the leaderboard for the same reason.
- The robustness slice is chosen where the audio is least like a neutral record of what happened. It records how often the system abstains, what the fallback does, and how much correction people end up carrying — the fall from 7.4 to 0.4 to 0.3 is the cost of that correction written down.
- Lifecycle evidence covers privacy, consent, incidents and downstream outcomes, none of which exist until after release. FCC 24-17 fixes the consent half as a legal precondition, and FCC 24-104's $6,000,000 forfeiture is what the incident half looks like when nobody wrote it down.
Accuracy, latency, abstention and consent belong in one document; any one of them on its own is a claim about a component, not about the product.
Key takeaways
- Treat what reaches the model as measurement rather than as sound: microphones, codecs, clocks, channel layouts, preprocessing and product choices all shaped it before the model saw anything.
- A high benchmark score on curated clips is evidence about curated clips. Five commercial systems averaged 0.35 word error rate for black speakers against 0.19 for white speakers on the same task, and ASVspoof 5's 32 attack algorithms significantly compromised baseline systems built without them.
- Start by naming the acoustic event and the observation process, and do not consider the design finished until the output is bound to a specific decision.
- Signal task, interaction task and decision system ask related but different questions, so evidence gathered to answer one of them does not transfer to the other two.
- A correct-looking transcript can still be the wrong thing to sign. 42.4% of physician-signed notes still contained errors after two stages of human review, which is why a release review has to ask what the recording could never have captured and who is absorbing the residual.
- Put task accuracy on representative acoustic slices in the same report as privacy, consent, incident and downstream outcome measures — FCC 24-17 makes prior express consent a precondition no accuracy figure can satisfy — and let the four decide the release together.