Skip to content
AI.info

Speech and audio

Speech Enhancement and Noise Suppression

Explain enhancement objectives, spectral masking, waveform models, perceptual losses, artifacts, and downstream evaluation.

By the end you can

Example

The denoiser made the transcript cleaner and the evidence less trustworthy

A compliance recorder applied aggressive enhancement to calls before archiving them. The output sounded clear, which was the point of installing it. What the archive no longer held was the recording. Quiet names had been altered. Some non-speech cues had disappeared. No later forensic review could separate what the microphone had captured from what the model had reconstructed.

The uncomfortable part is that a release review would have signed this off. Ask only about intelligibility and transcript change on realistic slices and the recorder passes. What was lost is not in the score. It is in the archive. Every claim this lesson makes about that failure has a published counterpart, with a name, a date and a number attached. The rest of the lesson is those counterparts.

  • That is what this lesson is about: what an enhancement system is actually for, and how spectral masking, waveform models, perceptual losses, artifacts and downstream evaluation each shape what comes out of it.
  • One failure produced the whole result: a perceptual metric optimized until it hid word deletion.
  • The evidence anyone would have thought to ask for is intelligibility and ASR change on realistic slices. That is precisely the evidence that would have let this recorder through.
  • The practical response is to stop asking one system to answer three demands at once. Separate what has to be listened to, what has to be recognized, and what has to be archived.

Comparison

How much the denoiser starts with

A single-channel denoiser has one mixture and nothing else. Whatever it cannot take from the recording it has to take from learned or statistical structure about how sources behave. Multi-channel enhancement has more than one observation of the same scene. A downstream front end is not obliged to produce listenable audio at all, only representations good enough for the task sitting behind it.

These are three different questions wearing one name. What clears one of them clears neither of the others. A result quoted from one setting says nothing about the other two. The named experiments in the rest of this lesson are almost all single-channel experiments. That is the hardest of the three cases. It is also the one the compliance recorder was in.

FigureComparison · 3 columns

Single-channel denoising

Uses one mixture and relies on learned or statistical source structure.

  • Decision focus: Define the target
  • Useful evidence: Intelligibility and ASR change on realistic slices
  • Watch for: Optimizing a perceptual metric that hides word deletion
  • Best used when its assumptions are documented for speech enhancement and noise suppression

Multi-channel enhancement

Adds spatial cues from synchronized microphones.

  • Decision focus: Specify mixture conditions
  • Useful evidence: Signal distortion and speaker-preservation measures
  • Watch for: Training on noise mixtures unlike production acoustics
  • Best used when its assumptions are documented for speech enhancement and noise suppression

Downstream front end

Optimizes a recognizer or detector, which may not produce human-faithful audio.

  • Decision focus: Select model and loss
  • Useful evidence: Human preference under controlled playback
  • Watch for: Using enhanced audio as the sole legal or clinical record
  • Best used when its assumptions are documented for speech enhancement and noise suppression

The mixture does not determine the answer

Speech enhancement estimates a cleaner target from a noisy recording. Wherever noise and speech overlap in the same time–frequency region, the mixture does not determine the answer. The task is underdetermined there. Something other than the observation has to fill the gap. That is why none of the three can simply recover the original.

Systems fill it differently. One predicts a mask, another a spectrum, another a waveform, another a residual. Those look like architectural choices and mostly are not. A team should weigh single-channel denoising, multi-channel enhancement and a downstream front end against each other first. The architecture follows from that decision rather than driving it.

One property is worth holding on to for the rest of the lesson. Enhanced audio is an estimate, not a recovered original. An estimate can be wrong while sounding entirely plausible: a word deleted, musical noise, phase distortion, a voice that is no longer quite the speaker's, or synthetic speech-like components sitting where the speech used to be.

Masks, spectra, waveforms and residuals differ mostly in how each fills the region where speech and noise overlap. Picking one is picking which guess you are prepared to defend.

Case

PESQ was deleted from the catalogue in January 2024

A release review will want a number, and the field has one it reaches for by reflex. Most enhancement papers report PESQ. The body that wrote PESQ has withdrawn it. ITU-T P.862, the PESQ recommendation, is marked out of date. The whole P.862 family was deleted from the catalogue on 5 January 2024. The ITU points readers to P.863 instead.

None of that stopped anything. PESQ “remains widely used despite its withdrawal by the International Telecommunication Union (ITU)” — that is the opening of a 2025 survey from Fraunhofer IIS and the International Audio Laboratories Erlangen. Its authors document the fact rather than dispute it. They track which versions and which open implementations are current, because people are going to keep computing the score either way. A number can be standard practice and unsupported at the same time. A score quoted from it inherits both facts.

Position

What a denoiser sounds like is not what it kept

PESQ was struck from the catalogue and kept in the papers. That is worth saying plainly rather than filing as a footnote, because a score inherits the standing of the standard behind it. But the withdrawal is the smaller problem. The larger one is what a listening score is being asked to stand for in the first place. That question was answered by experiment nearly two decades ago.

Eight single-microphone enhancement algorithms — spectral subtractive, subspace, statistical-model-based and Wiener-type — were put in front of normal-hearing listeners. The material was IEEE sentences and consonants corrupted by babble, car, street and train noise at 0 and 5 dB SNR. Hu and Loizou published the result in the Journal of the Acoustical Society of America in September 2007: “With the exception of a single noise condition, no algorithm produced significant improvements in speech intelligibility.” Eight algorithms, four noise types, two SNRs. One condition out of all of that showed a significant gain in what the listener could actually make out.

The second finding is the one that should end the reflex of quoting a quality score. Hu and Loizou report that the algorithms previously rated best for overall quality were not the ones best at preserving intelligibility. The ranking by how good the audio sounds and the ranking by how much of the speech survives are not the same ranking. Replacing P.862 with P.863 does not touch this. A better listening score is still a listening score. It measures how the audio lands on the ear. The open question is what the audio still contains.

So the demo is not the evidence. Ask instead for intelligibility, and for the change in recognition error on realistic slices. Those two measure content rather than impression, which is why they belong in every review. The recorder is the reminder that they are a floor rather than a ceiling. It would have passed on both.

Hu and Loizou found the algorithms rated best for overall quality were not the ones best at preserving intelligibility. Perceptual plausibility can conceal a deleted word, and a listening score will not.

Visual

Keep the original recording — two forensic bodies already require it

Three things have to happen in order. Specify the mixture conditions the system will actually meet. Define the target signal it is supposed to recover. Preserve auditability, so the untouched recording can still be inspected afterwards. The recorder never met the third.

That third step is not a principle this lesson invented. It is written down twice. ENFSI's forensic speech and audio group states it as a requirement in its Best Practice Manual for Digital Audio Authenticity Analysis, approved in December 2022. Section 8.2, “In the Laboratory”, reads: “Practitioners should always create and work on a working copy of the audio evidence, and not on the original submitted copy.” The same section requires the working copy to be hash-verified with SHA-2 or SHA-3. Access to the evidence storage has to be protected by write blockers.

The second body says it independently, and says it about enhancement specifically. SWGDE's Best Practices for Enhancement of Digital Audio, released in August 2025, requires every examination to be carried out on a working copy. Section 10.1 requires the processes, settings and applications used to be documented “in sufficient detail to repeat or reproduce the final results”, with their versions and the time segments processed. That revision's history records Appendix A being revised for the inclusion of “deep learning filters (trained neural networks)”. The forensic community has already written the neural denoiser into its documentation requirement.

Under time pressure it is the first step that goes. Nobody writes the mixture conditions down, so the definition of the target is never tested against them. It does not stop being used for that. It travels down the entire chain to the moment auditability matters, and by then it has been carried so long that nobody remembers it was ever a choice. The compliance recorder did not fall short of an ideal. It archived the enhanced audio in place of the original, which two published standards forbid.

FigureHierarchy · 4 levels
  • Define the target

    Choose intelligibility, ASR robustness, listening comfort, or signal fidelity as the primary outcome.

    • Specify mixture conditions

      Document noise types, SNR, reverberation, devices, and nonstationarity.

      • Select model and loss

        Balance spectral, waveform, perceptual, and downstream objectives.

        • Preserve auditability

          Retain original audio, model version, parameters, and confidence or quality indicators.

By the time auditability matters, the target you defined has stopped being a choice and become an untested assumption carried through the whole chain.

Example

Four products, four tolerances for suppression

Part of the recorder's trouble is that enhancement is not one product. It is deployed in at least four places that disagree about how much suppression is acceptable. Each of them pairs its own unit with its own thing to demonstrate.

The ASR case has been taken apart in detail, and the mechanism is not the obvious one. Researchers at NTT and Doshisha University split the error of a single-channel enhancement system into two parts. One is the noise it failed to remove. The other is an “artifact” component: the part of the error that cannot be written as a linear combination of the speech and noise sources. They separated the two by orthogonal projection-based decomposition. Their 2022 finding is blunt: “We experimentally identify the artifact component as the main cause of performance degradation, and we find that mitigating the artifact can greatly improve ASR performance.” The remedy is almost comic in its directness. Iwamoto and colleagues added a scaled copy of the observed signal back into the enhanced signal. The signal-to-artifact ratio rose, and ASR improved on both simulated and real recordings. What breaks the recognizer is not the noise the denoiser left behind. It is what the denoiser added.

  • Teleconferencing wants steady noise gone, and it wants interruptions and speaker cues left intact, because those are how a conversation actually works.
  • For ASR, enhancement is not automatically an improvement. Iwamoto and colleagues traced the degradation to the artifact component the enhancement itself introduces, and recovered ASR performance by mixing the observed signal back in.
  • Hearing devices cannot suppress aggressively at all: latency, comfort and spatial awareness each constrain what the wearer can be asked to accept.
  • Forensics is where the compliance recorder belongs, and its demand is the strictest of the four. ENFSI and SWGDE both require work on a copy and documented provenance for every processing step.

Example

Musical noise is a kind name for damage

Forensic provenance depends on describing what happened to the audio accurately. Two of the words in common use sound like siblings and do very different jobs. Musical noise is an artifact with an agreeable name. Speech distortion is damage to the very thing the system was installed to rescue.

A time–frequency mask and a perceptual loss are the tools that produce both, which is how the confusion starts. Call speech distortion "musical noise" in a report and you have described damage as an artifact. Different proof, different unit, and often a different person deciding what to do about it. The artifact decomposition is the precise version of this distinction. It separates what the enhancement failed to remove from what the enhancement itself put there. Only the second one is the artifact component.

  • A time–frequency mask is a weighting applied across time and frequency to emphasize an estimated source.
  • Musical noise is the tonal artifact produced by irregular residual spectral components.
  • A perceptual loss is an objective designed to correlate with human judgments or with learned representations.
  • Speech distortion is any change that enhancement introduces to the desired speech component itself.

Example

Even the official listening test refuses to answer with one number

Naming the damage correctly is what lets an evaluation catch it. The standards bodies and the large public challenges have both already concluded that one number will not do it.

The ITU has a recommendation written specifically for judging noise suppressors, and its design is the argument. ITU-T P.835 has been in force since 2003 and was last revised in July 2026. It does not ask listeners for a quality score. It asks them for three: the speech signal (SIG), the background noise (BAK) and overall quality (OVRL), rated separately. Microsoft researchers put it plainly — the “ITU-T Rec. P.835 subjective evaluation framework gives the standalone quality scores of speech and background noise in addition to the overall quality.” The official listening test was built on the assumption that a suppressor can clean the background while damaging the speech, and that one number would hide it.

The URGENT 2024 Speech Enhancement Challenge reached the same place from the other direction, at scale. Its 21 finalist teams were scored on 13 evaluation metrics across five categories: non-intrusive SE metrics, intrusive SE metrics, downstream-task-independent metrics, downstream-task-dependent metrics, and the subjective metric MOS. The MOS component ran as ITU-T P.808 tests on a 300-sample subset of the blind test data, half simulated and half real-recorded. Eight vetted listeners rated each sample. The organisers' own conclusion, published in 2025: “On the other hand, using solely a single category of metrics, especially non-intrusive metrics, for SE performance assessment can be unreliable or even misleading.” Combining the families tracked human judgement. Leaning on one did not.

So when intelligibility and ASR change disagree with artifact and original-versus-enhanced review, that conflict is not noise in the process. It is what the evaluation is run to produce. It should mark a decision point. The moment a word goes missing while the audio improves, the system is obliged to keep the original and fall back to it.

  • The core evidence for the task itself is intelligibility and ASR change on realistic slices.
  • Beside it belong measures of system behavior — signal distortion, and whether the speaker still sounds like the speaker, which is what P.835 separates as SIG from BAK.
  • The robustness slice is human preference under controlled playback, and P.835 and P.808 are the named methodologies for running it rather than improvising it.
  • Across the life of the system, run artifact and hallucination review, and keep a standing comparison of the original recording against the enhanced one.

P.835 splits the listener's judgement into SIG, BAK and OVRL, and URGENT 2024 needed 13 metrics across five categories. Both were built because one score is unreliable or even misleading on its own.

Steps

Build an enhancement evidence matrix

An evidence matrix is how that comparison stops depending on whoever happens to be listening. It has one job: to let another team challenge a score that improved while a word went missing. That is the challenge nobody brought against the compliance recorder.

What specifying mixture conditions actually costs is on the record. The INTERSPEECH 2020 Deep Noise Suppression Challenge, run by Microsoft, built its training speech out of LibriVox chapters rated with ITU-T P.808. It kept the upper quartile: chapters with 4.3 ≤ MOS ≤ 5. That came to 500 hours of speech from 2150 speakers. The noise side was about 150 audio classes and 60,000 clips drawn from AudioSet and Freesound, augmented with a further 10,000 clips from Freesound and DEMAND. That is what the second step of this process looks like when someone does it properly rather than naming it.

The test set is the part worth copying. The organisers released four categories of 300 clips each: synthetic without reverb, synthetic with reverb, real recordings collected internally at Microsoft, and real recordings from AudioSet. The reason is in their own abstract — “However, often the model performance degrades significantly on real recordings.” A synthetic test set drawn from the training distribution flatters the model. The split exists so the gap is visible instead of averaged away.

Three things then go on your page. Write down what defining the target signal assumes, since that is the assumption shown earlier travelling untested down the chain. Add one counterexample, a case where the definition would mislead you. Then say what preserving auditability triggers when the two disagree: who is told, what is kept, and which version becomes the record.

FigureProcess · 4 steps
  1. 1. Choose three purposes

    Separate listening, recognition, and archival requirements.

  2. 2. Create controlled mixtures

    Vary SNR, noise stationarity, reverberation, and speaker profile.

  3. 3. Compare outputs and originals

    Audit words, identity cues, timing, and non-speech events.

  4. 4. Write a release boundary

    State when enhancement may assist and when original evidence must remain primary.

Your matrix earns its keep the moment it names a case where the cleaned-up version helps and a case where the untouched recording stays the record. Like the Deep Noise Suppression challenge test set, it keeps the real recordings scored separately from the synthetic ones.

Key idea

A word can disappear and the score improve

Enhancement stays inside its boundary until one of four conditions is violated. The compliance recorder is recognizable in more than one of them.

The first is optimizing a perceptual metric that hides word deletion. That is the failure this lesson has been about from the opening, and the one Hu and Loizou measured. The second is training on noise mixtures unlike production acoustics. Those are the mixture conditions nobody specified, arriving later as the degradation on real recordings that the Deep Noise Suppression challenge built a separate test category to expose. The third is using enhanced audio as the sole legal or clinical record. That is exactly what an archive of enhanced calls is, and exactly what ENFSI and SWGDE forbid. The fourth is amplifying bias against quiet, accented, child or impaired speech. It should make those altered quiet names uncomfortable reading. The speech a system is least certain about is the speech it is most willing to rewrite.

That fourth condition is no longer hypothetical. The second URGENT Speech Enhancement Challenge, run in 2025, received 32 submissions. The winning system was discriminative, and most other competitive entries were hybrid. Saijo, Zhang and their co-organisers report: “Analysis reveals some key findings: (i) some generative or hybrid approaches are preferred in subjective evaluations over the top discriminative model, and (ii) purely generative SE models can exhibit language dependency.” Both halves belong here. A system that generates its output can behave differently depending on the language it is given. And the systems listeners preferred were not the system that won. Perceptual preference and measured performance came apart in a controlled public challenge, exactly as they did for Hu and Loizou in 2007.

All four conditions end in the same place. Each of them lets a wrong estimate through by checking how clean the output sounds instead of checking what survived inside it.

Perceptual scores can rise while words vanish, so a claim resting on that number alone describes how the audio sounds, not what it still says.

Key takeaways