Skip to content
AI.info

Speech and audio

Sampling, Aliasing, Quantization, and Bit Depth

Explain sample rate, Nyquist limits, anti-alias filtering, quantization, bit depth, dither, and clock quality.

By the end you can

Comparison

Aliasing, rounding, clipping: three different damages

The same conversion can go wrong in three ways. A recording can suffer any one of them without showing the other two.

Aliasing is high-frequency content turning up at the wrong lower frequencies after inadequate sampling. Energy the chain could not represent does not disappear. It comes back wearing the wrong frequency. Quantization error is what is left over when a continuous amplitude is rounded to one of a finite set of levels. Clipping is what happens when the signal runs past full scale and the numbers stop going up. Three problems, one conversion chain, three different tests. Passing the test for one says nothing about the other two.

What they have in common is a habit of hiding behind a single number. The sample rate looks like a summary of the conversion. It is not even a complete summary of the first damage. A wildlife team discovered this on a season of recordings it could not fix.

FigureComparison · 3 columns

Aliasing

High-frequency content appears at incorrect lower frequencies after inadequate sampling.

  • Decision focus: Set the required bandwidth
  • Useful evidence: Usable passband and stop-band attenuation
  • Watch for: Using the Nyquist frequency as the entire practical passband
  • Best used when its assumptions are documented for sampling, aliasing, quantization, and bit depth

Quantization error

Amplitude is rounded to one of a finite set of representable levels.

  • Decision focus: Choose sample rate and clock
  • Useful evidence: Effective number of bits and measured noise floor
  • Watch for: Upsampling and describing the result as higher-fidelity evidence
  • Best used when its assumptions are documented for sampling, aliasing, quantization, and bit depth

Clipping

Values beyond the available range are truncated, causing nonlinear distortion.

  • Decision focus: Choose amplitude resolution
  • Useful evidence: Clock drift and inter-channel synchronization error
  • Watch for: Recording too close to full scale and clipping unpredictable transients
  • Best used when its assumptions are documented for sampling, aliasing, quantization, and bit depth

Example

The bat calls were never in the file

A wildlife team upsampled recordings made at 16 kHz, expecting to recover bat calls. The new file had more samples than the original. It had no bat calls. The microphone had never captured those frequencies in the first place.

How far short 16 kHz fell can be stated exactly, because the band bats occupy has been measured. A 2017 paper in Scientific Reports opens with the range: "Across vespertilionids (the most species rich bat family, ~420 species) echolocation call PFs range from 10 to 150 kHz, and mass from 2 to 60 grams." Most FM calls peak between 20 and 60 kHz. The extreme in the same paper is the trident bat Cloeotis percivali, printed there as "Cleotis percivalis", with a 212 kHz carrier — the highest frequency pure tone documented from the natural world. Behind those figures sit 86 species and 260 specimens with mass and gape data, plus 69 species and 200 specimens with forearm data.

Set those numbers against the recording. A 16 kHz sample rate puts the Nyquist ceiling at 8 kHz. That ceiling sits below even the lowest vespertilionid peak frequency of 10 kHz. It sits roughly nineteen times below the 212 kHz carrier. The capture did not lose part of the call band. It never reached it.

The arithmetic behind the team's expectation was not wrong. Half of 16 kHz is 8 kHz, and content above that cannot be represented uniquely. The mistake was reading that limit backwards, as though everything beneath it were guaranteed to be there. Nobody asked which frequencies the recording chain had ever passed. Half the sample rate had been treated as the answer.

  • The team paid for a lesson that covers a whole chain of choices — sample rate, Nyquist limits, anti-alias filtering, quantization, bit depth, dither and clock quality. Settle any one of them in isolation and the rest go unasked.
  • Underneath it was one failure: using the Nyquist frequency as the entire practical passband. That turns a ceiling into a promise. The promise was 8 kHz, against a call band published as 10 to 150 kHz.
  • What nobody had measured was the usable passband and the stop-band attenuation. Which frequencies actually survived the chain, and how hard everything beyond them was pushed down.
  • The practical response costs nothing and has to happen before recording. Name the expected bandwidth of the event you are after, and its shortest timing feature. For vespertilionids that figure was in print, in Scientific Reports, before the season began.

Upsampling adds numbers, not information

That upsampled file is worth dwelling on. It is the cleanest demonstration there is of which parts of a conversion can be undone afterwards. None of them.

Sampling records a continuous-time signal as discrete values taken at fixed instants, and two things can go wrong at that instant. Frequencies above the usable band can fold down into lower ones unless they are filtered before conversion. Quantization, which maps continuous amplitudes to finite levels, introduces a different class of error entirely. Both happen once, at the moment the file is written.

Which is why upsampling later is an arrangement rather than a recovery. It can be well worth doing — a downstream stage may need everything at a common rate — but it cannot bring back what the sensor, the analog path, the anti-alias filter or the original sample rate removed. Bit depth behaves the same way. More bits do not repair clipping, poor calibration or a noisy microphone. They write that damage down more finely.

The converse has been measured too. Take high-resolution playback, insert a 16-bit/44.1 kHz analogue-to-digital-to-analogue loop into it, and ask professional engineers, recording students and audiophiles to say when the loop is there. Meyer and Moran ran that test double-blind for a year and published it in 2007. Their result: "The test results for the detectability of the 16/44.1 loop on SACD/DVD-A playback were the same as chance: 49.82%. There were 554 trials and 276 correct answers." Extra rate and depth downstream of what the chain delivered bought, in 554 trials, nothing a listener could name.

So a recording has two separate questions to answer, and each needs its own numbers. Usable passband and stop-band attenuation ask what survived the conversion. Effective number of bits and measured noise floor ask how finely what survived was written down. Neither pair should stand in for the complete release decision.

Extra bits and a higher rate cannot add what a clipped capture never held: 276 correct answers in 554 trials, 49.82% — chance.

Case

Why telephony still throws away everything above 4 kHz

The wildlife team lost a band without deciding to. Telephony lost a wider one on purpose, wrote the decision down, and has lived inside it ever since. That makes it the clearest available record of what a sample rate actually commits you to.

ITU-T G.711 has been in force since 1988. It codes voice at eight bits a sample and 64 kbit/s, and the RTP profile carries it as PCMU and PCMA at a clock rate of 8,000 Hz. That clock is the whole story. An 8 kHz sample rate keeps nothing above 4 kHz, which is where much fricative energy sits. What the standard discards is not a decorative top end. It is part of how one consonant is told from another. Every telephone call has been making the wildlife team's trade since before the team recorded anything. The difference is that G.711 knew it was making it.

Wideband G.722 doubles the sampling rate to 16,000 Hz and hears a different signal. Its title promises 7 kHz audio-coding within 64 kbit/s. The scope clause is more careful: "This Recommendation describes the characteristics of an audio (50 to 7 000 Hz) coding system which may be used for a variety of higher quality speech applications." Sixteen thousand samples a second put the Nyquist ceiling at 8 kHz. The delivered band is 50 to 7 000 Hz. The gap between those two figures is not sloppiness. The standard lists what stands in it: an input anti-aliasing filter, a sampling device operating at 16 kHz, and an analogue-to-uniform digital converter with 14 bits. The filter is a required functional unit ahead of the sampling device, and a filter that does its job takes its cut below the ceiling.

Note where doubling lands G.722. 16,000 Hz is exactly the rate the wildlife team was already recording at. Telephony fixed its sampling rate decades ago, and much production audio still lives there. A system asked to work on that audio inherits a passband somebody else settled long before the question was asked. It inherits the slack too: 1 000 Hz between ceiling and delivered band. The standard writes that down. A file header does not.

Key idea

Nyquist is not the usable passband

G.711 can be honest about its ceiling because the ceiling is written into the standard. Trouble starts where the rate is inherited rather than chosen, and the number gets read as a promise about which frequencies a recording holds. Two of the four shortcuts below come from taking that promise literally. The other two come from reading bit depth the same way.

1) Using the Nyquist frequency as the entire practical passband — the wildlife team's move. 2) Upsampling and describing the result as higher-fidelity evidence, which is the same move committed a second time, after the loss. 3) Recording too close to full scale and clipping unpredictable transients. 4) Assuming nominal bit depth equals effective dynamic range.

The first shortcut has a documented counter-example in the narrowband case, not just the wideband one. ITU-T G.712, approved in 2001, sets its in-band requirements for the PCM telephone channel over 300 Hz to 3400 Hz, and states its return loss requirement over that same range. The companion Recommendation G.711 samples at 8 kHz, which permits 4 kHz. So the arithmetic ceiling is 4 kHz and the specified channel stops at 3400 Hz. The region between them is not spare capacity. G.712 reserves 3400–4600 Hz for the anti-aliasing filter and gives it its own attenuation template, Figure 10/G.712. A note in the standard says why: "NOTE 2 – Attention is drawn to the importance of the attenuation characteristic in the range 3400 Hz to 4600 Hz." Eight thousand samples a second, a 4 kHz Nyquist frequency, 300–3400 Hz actually delivered, and a written template governing the band in between.

All four shortcuts are hard to catch because each produces a file that looks better by its own measure: more samples, more bits, a level sitting confidently near the top of the scale. Every one of those numbers is moving in the direction that usually signals quality. But the sample rate marks only a ceiling. The filters, the transducer and the analog path all take their cut below it. G.711 tells you where its ceiling is. G.712 tells you how far under it the channel actually delivers. A file header does neither.

The sample rate marks a ceiling, not a delivered band: 8 kHz sampling permits 4 kHz, and G.712 specifies the channel over 300 Hz to 3400 Hz.

Example

Passband and Nyquist frequency are not synonyms

Half of those shortcuts turn on one confusion between two words. Nyquist frequency and usable passband are the pair most often swapped. Swapping them changes what has to be proved, the unit it is proved in, and who may sign the result off. Sample rate, Nyquist frequency, quantization and dither each answer a different question about the same conversion. Anything written about that conversion has to use the four terms exactly.

  • Sample rate is the number of time samples recorded per second — 16 kHz in the wildlife recordings, 8,000 Hz in a G.711 call, 16 kHz into the 14-bit converter G.722 specifies.
  • Nyquist frequency is half the sample rate. It rests on Theorem 1 of Shannon's 1949 paper, and the theorem is a conditional: "If a function f(t) contains no frequencies higher than W cps, it is completely determined by giving its ordinates at a series of points spaced 1/2W seconds apart." Read the sentence as written. It fixes what must be true above W. It claims nothing about what a recording holds below it.
  • Quantization is the mapping of an amplitude to one of a finite set of digital levels — the point at which amplitudes stop being continuous and start being rounded.
  • Dither is noise added deliberately, to decorrelate some of the artifacts quantization leaves behind. The standard recommendation is specific, and a 1992 survey by Lipshitz and two colleagues states it: "We recommend the use of nonsubtractive, triangular-pdf dither of 2-LSB peak-to-peak amplitude for most audio applications requiring multibit quantization or requantization operations, since this type of dither renders the first and second moments of the error signal constant with respect to the input, while incurring the minimum increase in error variance." They call that choice unique and optimal: it renders the first and second moments of the total error input independent, while minimizing the second moment. It is not free. Tripling the noise spectral density costs a 4.8 dB rise in the noise floor relative to subtractive dither.

Visual

The clock is the step teams skip

These confusions get their opening at particular points in the work, and it is worth seeing where. An acquisition path has an early step where somebody states the bandwidth the application requires, and a much later one where the complete capture path is validated against it. Between the two sits the step teams skip: choosing the sample rate and the clock deliberately, as a decision that could have come out differently.

Skip it and the bandwidth fixed at the start is never questioned again. By the time the complete-path validation runs, that figure is no longer an input under test. It has become the standard the test is measured against. The wildlife team's 16 kHz sat in exactly that position. Nothing they did afterwards was ever going to disagree with it. No validation run against an 8 kHz ceiling can report the absence of a 10 to 150 kHz call band it was never asked about.

The standards show the same three steps done in the open. G.722 fixes the required band first, 50 to 7 000 Hz. Then it names the clock and the converter, 16 kHz sampling and 14 bits. Then it writes the filter template the delivered path has to meet. The order is the whole of the discipline.

FigureLayers · 4 layers
  1. 01

    Set the required bandwidth

    Start from the highest meaningful frequency and the transition band of the acquisition filter.

  2. 02

    Choose sample rate and clock

    Provide margin for filtering and verify timing stability across channels.

  3. 03

    Choose amplitude resolution

    Balance bit depth, headroom, noise floor, storage, and downstream computation.

  4. 04

    Validate the complete path

    Test real sensors and converters rather than relying on nominal file metadata.

Bandwidth you fixed early is baked into the path you later declare valid, so a wrong figure passes the check instead of failing it.

Steps

Design an acquisition specification

The repair for a skipped step is a document another team can argue with. An acquisition specification is worth writing only if someone else can use it to challenge the claim that cost the bats: that a sample rate on its own tells you which band you actually get.

Three lines are enough. Write down what your choice of required bandwidth assumes. Write down one case where that assumption fails. Write down what validating the whole capture path does about it. If the second line is hard to write, the first has not yet been made specific enough to be wrong.

G.722 is a model of the three lines. Its assumption is stated as a delivered band, 50 to 7 000 Hz. The case where a bare ceiling would fail is the 1 000 Hz between that band and the 8 kHz Nyquist frequency. What it does about it is a named functional unit, the input anti-aliasing filter ahead of the sampling device. Write your own three lines with figures of that kind, and a reviewer can disagree with them.

FigureProcess · 4 steps
  1. 1. Choose a target event

    Name its expected bandwidth and shortest timing feature.

  2. 2. Allocate margins

    Reserve transition band, headroom, and synchronization tolerance.

  3. 3. Estimate data volume

    Calculate channel count, sample rate, bit depth, and retention.

  4. 4. Plan a validation recording

    Include tones, impulses, silence, and realistic environmental noise.

A specification never exercised on tones, impulses, silence, and real room noise is a statement of intent, not a measurement of the capture path.

Example

Where bandwidth and headroom pull apart

Exercise that specification and it returns numbers that do not all agree. That is the reason for collecting them. The two pairs pull apart under load. What the passband preserves can sit at odds with the clipped-sample rate and the headroom left for transients, and no single figure reconciles the two. An evaluation of the conversion chain earns its keep at exactly that point, where its own measurements point in opposite directions and somebody has to choose.

One of those measurements has a written test method. That matters, because the fourth shortcut — assuming nominal bit depth equals effective dynamic range — is refuted by the definition itself. IEEE Std 1241-2010 defines effective number of bits this way: "For an input sine wave of specified frequency and amplitude, after correction for gain and offset, the effective number of bits (ENOB) is the number of bits of an ideal ADC for which the rms quantization error is equal to the rms noise and distortion of the ADC under test." It is computed from the measured noise-and-distortion relative to full-scale range. Nominal bit depth is a number in a file header. ENOB is the result of putting a converter on a bench with a sine wave. Only one of them is evidence.

Four kinds of evidence make the choice possible. Alongside them, the portfolio should show what happens when the band a recording actually covers is too narrow for the question being asked — the wildlife case arriving as a live input rather than a post-mortem — and the system has to abstain or fall back.

  • Evidence about the core task is usable passband and stop-band attenuation. Their absence is the whole of the wildlife story: 8 kHz of ceiling against a published call band of 10 to 150 kHz.
  • Evidence about system behavior is the other pair, effective number of bits and measured noise floor — ENOB as IEEE Std 1241-2010 defines it, measured, not read off a datasheet line.
  • The robustness slice is about time rather than frequency: clock drift, and inter-channel synchronization error.
  • Lifecycle evidence is clipped-sample rate and transient headroom, which record how close the capture kept running to full scale.

Report usable passband and stop-band attenuation together with clipped-sample rate and transient headroom.

Key takeaways