Speech and audio
Dereverberation, Echo Cancellation, and Acoustic Feedback
Distinguish reverberation, acoustic echo, feedback, and their adaptive cancellation strategies in real systems.
By the end you can
- Define dereverberation, echo cancellation, and acoustic feedback as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish dereverberation, acoustic echo cancellation, and feedback control without treating them as interchangeable
- Trace the workflow from identify the phenomenon through control residual and failure
- Evaluate dereverberation, echo cancellation, and acoustic feedback using echo return loss enhancement and residual audibility and evidence from difficult deployment slices
Three problems that sound like one
Three things go wrong with sound in a room, and they arrive sounding alike. Reverberation is the room's extended response to a source. Echo is delayed playback or a reflection, and it may come with a reference signal — a copy of what was played, already sitting in memory. Acoustic feedback is a loop, and a loop can become unstable. Different structures leave you different things to measure. One generic suppression model is rarely enough.
Feedback alone has half a century of control theory behind it. The survey of it runs to 40 pages in Proceedings of the IEEE, published in 2011 by van Waterschoot and Moonen. Its abstract sets out the scope: “The acoustic feedback problem has intrigued researchers over the past five decades, and a multitude of solutions has been proposed. In this survey paper, we aim to provide an overview of the state-of-the-art in acoustic feedback control, to report results of a comparative evaluation with a selection of existing methods, and to cast a glance at the challenges for future research.” Fifty years of dedicated method development, and a comparative evaluation of the methods against each other. No field builds that for a special case of noise suppression.
Only one of the three hands you a reference, and echo cancellation is built on it. It assumes a usable reference, and a path that can be tracked. Someone moving the device, nonlinear loudspeakers, clock mismatch, double-talk and feedback each violate one half of that model. Which architecture you pick matters less than whether you named the problem correctly first. Suppression aimed at the wrong phenomenon is still suppression.
Mistake one of these three for another and the suppressor you ship will be tuned against a phenomenon that was never the problem.
Visual
Timing and correlation tell them apart
Two things do the naming: when the unwanted sound arrives, and whether it correlates with something you already hold. Playback echo is delayed, and it correlates with the digital reference that produced it. Reverberation is the room's own decay. It correlates with the source, but there is no clean copy of the source to subtract. Feedback returns the output to the input, so what it correlates with is the system's own past. That is what makes it a loop that can grow rather than a response that dies away.
That test has to come first, for a structural reason. The work runs in three stages: identify the phenomenon, map the available references, control residual and failure. They are ordered by dependency, not by convenience. Mapping the references consumes the answer the first stage produced. It can inherit the label, not question it. Only residual and failure control puts the label under strain, and it sits at the end. By then the suppressor has already been built around whatever name it was given.
1. Identify the phenomenon
Use timing, correlation, decay, and system topology to distinguish reverberation, echo, and feedback.
2. Map available references
Record far-end playback, loudspeaker signals, channel clocks, and microphone geometry.
3. Estimate or adapt the path
Track the room and device response while protecting near-end speech.
4. Control residual and failure
Add double-talk detection, residual suppression, stability limits, and user fallback.
Call feedback echo, and the reference mapping will accept the label without complaint; residual control is where the mistake finally surfaces.
Example
A video call denoiser could not cancel the loudspeaker
A meeting device made exactly that mistake. It treated far-end playback — the far end's audio, coming back out of its own loudspeaker — as ordinary background noise. Everything needed to name it correctly was present. The playback correlated with a known digital reference, and it changed with the room acoustics. Filed as noise, none of that was used. The denoiser left residual echo. During double-talk it sometimes removed near-end speech instead, cutting off the person in the room to quieten a signal it had misnamed.
Averaged echo-removal figures would have looked acceptable here. That is how a team ends up suppressing what is left over before the canceller has settled, with nobody reading the one number that would have shown it. The figures later in this lesson are the field's own measurements of that gap — between a removal average and a listener.
- The decision underneath this device is the one the whole lesson turns on. Tell reverberation, acoustic echo and feedback apart in a real system, then choose an adaptive cancellation strategy that fits the phenomenon you actually have.
- The failure that follows is almost always the same one — residual suppression switched on before the linear echo estimate is stable.
- The evidence that would have caught it was there to be read: echo return loss enhancement alongside residual audibility. How much echo went, and whether anyone could still hear what stayed. As the challenge organisers document below, the first of those two numbers is only defined for the quiet single-talk case the device passed.
- The response is unglamorous. Mark the source, the playback, the room paths, the microphones, the references and the feedback loops before touching a suppressor.
Case
From 2,500 rooms to 10,000 devices, and from 40 ms to 20 ms
The field reached the same conclusion by a longer road, and left a paper trail with numbers on it. Echo cancellation is old enough to have a network standard and new enough to have a run of challenges. ITU-T G.168, on digital network echo cancellers, is still in force in its 2015 edition. The challenge series went the other way, onto the hardware in the room, and it did so in deliberate steps.
The training and test data grew every year: 2,500 real environments for ICASSP 2021, then 5,000 for INTERSPEECH 2021, then 7,500 for ICASSP 2022, then 10,000 for ICASSP 2023. The permitted algorithmic plus buffering latency was cut from 40 ms to 20 ms for the 2023 edition. The first three challenges drew 49 participants. The 2023 edition received 20 entries, 17 non-personalized and 3 personalized. Cutler and nine colleagues tabulate the series in the ICASSP 2023 challenge report. Of the 2023 data the abstract says: “These datasets consist of recordings from more than 10,000 real audio devices and human speakers in real environments, as well as a synthetic dataset.”
The INTERSPEECH 2021 edition, the one at 5,000, described its data the same way — “recordings from more than 5,000 real audio devices and human speakers in real environments”. It ranked entries not by an objective figure but by listeners, with winners “selected based on the average Mean Opinion Score”. Both halves of that verdict would have caught the meeting device. Its trouble was its own loudspeaker in its own room, and what was wrong with the result was audible.
Comparison
With a reference, without one, in a loop
What separates the three problems is what each one gives you to aim at.
Dereverberation reduces late room energy with no perfect dry reference to aim at. That is not an inconvenience of practice. It is written into the benchmark. The REVERB challenge, from Kinoshita and eleven colleagues in 2013, fixes the conditions and then withholds the room: “SimData simulates 6 different reverberation conditions: 3 rooms with different volumes (small, medium and large), 2 types of distances between a speaker and a microphone array (near=50 cm and far=200 cm). The reverberation times (T60) of the small, medium, and large-size rooms are about 0.25, 0.5, 0.7 s, respectively.” Stationary noise is added at 20 dB SNR. Alongside the simulations sit real MC-WSJ-AV recordings, made in a meeting room with a T60 of 0.7 s. The rules explicitly forbid entrants the room's own parameters — the reverberation time, the direct-to-reverberation ratio, the impulse responses themselves. The dry reference is absent by design. A method that quietly leans on it is disqualified, not merely optimistic.
Acoustic echo cancellation and feedback control work from other information, and each has its own literature to be measured against. The feedback survey does not only describe the methods; van Waterschoot and Moonen report “results of a comparative evaluation with a selection of existing methods”. That is what it means for a problem to be theorised in its own right. The meeting device is what happens when the three are not diagnosed in their own terms. It held the most usable information of the three, a digital copy of the offending signal. It spent it as though it had none.
Dereverberation
Reduces late room energy without a perfect dry reference.
- Decision focus: Identify the phenomenon
- Useful evidence: Echo return loss enhancement and residual audibility
- Watch for: Applying residual suppression before the linear echo estimate is stable
- Best used when its assumptions are documented for dereverberation, echo cancellation, and acoustic feedback
Acoustic echo cancellation
Uses a known playback reference to estimate the echo path.
- Decision focus: Map available references
- Useful evidence: Near-end speech distortion during single- and double-talk
- Watch for: Mistaking near-end speech for echo during double-talk
- Best used when its assumptions are documented for dereverberation, echo cancellation, and acoustic feedback
Feedback control
Prevents or suppresses a closed loop that can grow unstable.
- Decision focus: Estimate or adapt the path
- Useful evidence: Decay reduction and intelligibility under reverberation
- Watch for: Ignoring sample-clock drift between playback and capture
- Best used when its assumptions are documented for dereverberation, echo cancellation, and acoustic feedback
Key idea
Double-talk is where the canceller gets it wrong
Once the label is right, the ordering is the next place to lose. Four shortcuts show up again and again in echo work. Every one is defensible on the day it is taken, and every one costs something later.
1) Applying residual suppression before the linear echo estimate is stable. 2) Mistaking near-end speech for echo during double-talk. 3) Ignoring sample-clock drift between playback and capture. 4) Training only in fixed rooms while the device moves in use.
Each shortcut bets against a different way the two assumptions break: a loudspeaker driven nonlinear, clocks drifting apart, a device picked up and carried, a feedback loop building, both ends talking at once. The meeting device took the first shortcut and paid for it through the second. That is the usual sequence.
That second cost is now the field's largest measured deficit. It is not the echo that survives — it is the talker who does not. The ICASSP 2023 challenge report gives the remaining headroom to a perfect score as 0.26 MOS for double-talk echo, 0.64 for “double talk other” (missing audio, distortions and cut-outs) and 0.74 for single-talk near-end speech quality. The top 2023 model, MS-1, scored 4.68 MOS on double-talk echo against only 4.26 on double-talk other. The organisers name the two remaining problems themselves: “The two largest areas for improvement are (1) Single Talk Near End quality, which is affected by background noise, reverberation, and capture device distortions, and (2) Double Talk Other Degradations, which includes missing audio, distortions, and cut-outs.” Echo removal in double-talk is very nearly solved. What the best systems still do is damage the near-end voice while doing it. That was the meeting device's own shortcut, and it still costs more than the fault it was meant to fix.
Residual suppression applied before the linear estimate settles buys quiet on day one and eats near-end speech later: 0.26 MOS of headroom left on double-talk echo, 0.64 on the damage done to the person in the room.
Analogy
Following footprints in a hall of mirrors
How much that reference is worth is easiest to see by taking it away. Echo cancellation predicts repeated footprints left by a traveler whose stride is already known. The meeting device had the stride written down and read the prints as weather. Dereverberation reads traces left by a room nobody has measured, and REVERB makes that literal by forbidding entrants the room's own parameters. But footprints are discrete and countable, where an acoustic path is a continuous filter. Real devices add nonlinearities and clock errors on top of it. The picture is friendlier than the problem it stands for.
Use the known reference when it exists; do not confuse residual echo with arbitrary noise.
Steps
Diagnose an acoustic loop
Diagnosing an acoustic loop is the exercise, and the meeting device is the shape of it. Write it up so another team can argue with the point at which you started suppressing residual echo, and with your evidence that the linear estimate had settled by then. That ordering is the claim most worth attacking, and the one most write-ups leave undefended. Record the states you tested in the terms the benchmarks use, so the write-up can be compared with something: far-end only, near-end only, double-talk, movement, silence. For the reverberant slice, state a T60 rather than the word “live”. Then set down what your diagnosis assumes. Give one counterexample that would break it. Say what you would change in the suppression once you know.
1. Draw the signal topology
Mark source, playback, room paths, microphones, references, and feedback loops.
2. Record controlled states
Capture far-end only, near-end only, double-talk, movement, and silence.
3. Inspect correlation and decay
Use impulse-response and reference alignment evidence.
4. Choose layered controls
Combine cancellation, suppression, gain control, placement, and product behavior.
A loop diagnosis that stops at the canceller is half a diagnosis; gain control, device placement and the product's own behavior sit on the same circuit.
Example
The removal figure correlated with listeners at 0.31
A diagnosis ends in numbers, and the meeting device is the reason no single one of them will do. Its removal average looked fine while the fault was plainly audible. The challenge organisers measured that gap rather than asserting it. On single-talk echo scenarios, ERLE correlated with subjective ITU-T P.808 ratings at PCC 0.31 and SRCC 0.23, against PCC 0.67 and SRCC 0.57 for PESQ. They also state where the figure is defined at all: “ERLE is only appropriate when measured in a quiet room with no background noise and only for single talk scenarios (not double talk), where we can use the processed microphone signal as an estimate for e(n).” A quiet room, single talk. That is the slice the meeting device passed, and the only slice on which its headline number was even meaningful.
Four kinds of evidence sit behind a release claim, and they answer four different questions. Report all four on the difficult slices as well as the clean ones. The slices where suppression ran ahead of the canceller are exactly the ones an average covers up.
- Echo return loss enhancement and residual audibility are the core pair, and the field's response to the first one failing was to build a second. AECMOS, from Purin and four colleagues in 2022, opens with the diagnosis: “Traditionally, the quality of acoustic echo cancellers is evaluated using intrusive speech quality assessment measures such as ERLE [1] and PESQ [2], or by carrying out subjective laboratory tests [3], [4]. Unfortunately, the former are not well correlated with human subjective measures, while the latter are time and resource consuming to carry out [5].” Their answer is a neural estimator of the subjective echo and other-degradation scores.
- Near-end speech distortion has to be measured during single-talk and again during double-talk. Double-talk is where the meeting device did its damage — and where 0.64 MOS of the field's remaining headroom sits, against 0.26 on the echo itself.
- Decay reduction and intelligibility under reverberation cover the harder slice, the one where the room rather than the loudspeaker is what the listener is fighting. REVERB is where those are scored, at T60 of about 0.25, 0.5 and 0.7 s and at 50 cm and 200 cm, with the room's parameters withheld.
- Convergence time, path-change recovery and feedback stability margin are the lifecycle numbers, and the last of them has a document behind it. The challenge organisers point to it: “ITU-T Rec. G.122 [17] defines AEC stability metrics, and ITU-T Rec. G.131 [18] provides a useful relationship of acceptable Talker Echo Loudness Rating and one-way delay time.” G.122 covers stability and talker echo in international connections, G.131 talker echo and its control. The ITU lists both as in force.
ERLE tracked listeners at PCC 0.31 where PESQ managed 0.67 — pair the removal figure with residual audibility, convergence time, path-change recovery and feedback stability margin, or report a number that is only valid in a quiet room.
Key takeaways
- Three phenomena, three structures — a room's extended response, a delayed playback that may arrive with a reference, and a loop that can become unstable. Each has its own literature. Feedback control alone fills a 40-page survey in Proceedings of the IEEE, published in 2011.
- Echo cancellation holds only while its two assumptions hold: a usable reference, and a path that can be tracked. Name what could break them on your hardware — a nonlinear loudspeaker, drifting clocks, a hand picking the device up.
- Naming the phenomenon comes first. The available references depend on that answer, and residual control does not test it until the end, when the suppressor is already built.
- Dereverberation is defined by what it is denied. REVERB withholds the room's own parameters and scores six conditions at T60 of about 0.25, 0.5 and 0.7 s, near=50 cm and far=200 cm. A method that leans on a dry reference is disqualified, not merely optimistic.
- Residual suppression applied before the linear echo estimate has converged hides an unstable path behind a quiet output. The collateral damage is now the larger deficit: 0.64 MOS of headroom on double-talk other degradations against 0.26 on double-talk echo, with MS-1 at 4.26 and 4.68 respectively.
- Pair the removal evidence with the lifecycle evidence, because the removal figure does not track listeners. ERLE correlated at PCC 0.31 and SRCC 0.23 with ITU-T P.808 ratings, against 0.67 and 0.57 for PESQ. It is only defined for a quiet room in single talk.