Speech and audio
Microphones, Channels, Gain, Clipping, and Calibration
Connect microphone response, placement, channel geometry, gain staging, clipping, calibration, and metadata to downstream model behavior.
By the end you can
- Define microphones, channels, gain, clipping, and calibration as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish absolute calibration, relative normalization, and domain normalization without treating them as interchangeable
- Trace the workflow from characterize the transducer through preserve calibration metadata
- Evaluate microphones, channels, gain, clipping, and calibration using frequency-response deviation and self-noise and evidence from difficult deployment slices
Key idea
The corpus whose channels had to be resynchronised two years later
Recordings stop supporting the claim made from them in four quiet ways. All four tend to go unnoticed until something downstream breaks. Device identity leaks into the target label when the classes and the recorders line up. Automatic gain control moves the level underneath the signal and changes the temporal dynamics. Channels get swapped, or the microphones of an array run unsynchronized. And the microphone and firmware metadata are discarded after ingestion, so none of the first three can be checked afterwards.
The third one is not a cautionary hypothetical. It has a size and a date. The CHiME-5 dinner-party corpus, released in 2018, recorded 20 parties in 20 homes. Six Kinect arrays of 4 sample-synchronised microphones each, plus binaural pairs worn by the 4 participants. The challenge's own data page does the arithmetic: “the total number of microphones per session is 32 (2 x 4 + 4 x 6)”. That is 50 h 12 min of audio in all — 40:33 train, 4:27 dev, 5:12 eval. The arrays were not reliably synchronised with each other. The Kinect recordings suffer audio frame-dropping, and the devices clock-drift relative to one another.
Two years later the same recordings came out again as CHiME-6. “Speech material is the same as the previous CHiME-5 recordings except for accurate array synchronization,” says the abstract of the 2020 overview paper. The baseline shipped a synchronisation tool that “compensates for two separate issues: audio frame-dropping (which affects the Kinect devices only) and clock-drift”. Synchronisation was not the only reason for the reissue. CHiME-6 also introduced a new Track 2 for unsegmented multispeaker recognition and diarization. But the repair was worth re-releasing 50 hours of audio that were already public.
The standard answer to all four modes is normalization, and normalization can align average scale. It cannot undo saturation. It cannot recover blocked frequencies, remove all device response, or reconstruct a source that was masked by reverberation and competing sound.
A 32-microphone, 50-hour public corpus had to be reissued with accurate array synchronization; normalization aligns scale and would never have touched it.
The microphone is part of the decision
Normalization stops where it does for a reason of order. What it rescales is the recorded number, not the capture that produced it. By the time a scaling factor is applied, saturation has already flattened the peaks. The microphone has already blocked the frequencies it does not pass, and its response is already stamped into the spectrum. Reverberation and competing sound have already masked the source. Scaling arrives after all of that and moves the whole mixture together. Frame-dropping and clock-drift are damage to the alignment between channels, and no per-file scaling factor addresses alignment at all.
So the microphone is not a neutral opening to the pipeline. It is a directional, frequency-dependent transducer embedded in a room and an electronic chain. Placement, polar pattern, gain control, clipping, channel synchronization and device metadata can all become part of the learned decision, whether that was intended or not.
That is why absolute calibration, relative normalization and domain normalization belong side by side, before anyone reaches for a different architecture. They buy different things. A fix aimed at the wrong one leaves the capture exactly as it was.
Whatever the room and the hardware added is already in the data, and the model will treat it as signal.
Example
ROC-AUC 0.846 before matching, 0.619 after
A respiratory-audio classifier can look like a screening tool while hearing something other than the disease. The clearest published instance analysed audio from 67,842 individuals, 23,514 of them PCR-positive, recruited through NHS Test-and-Trace and the REACT survey. Not a pilot, and not a small one. Coppock and colleagues reported it in Nature Machine Intelligence in February 2024: audio-based classifiers showed no evidence of improving on a simple symptoms checker.
Unadjusted, the classifiers looked like a screening tool: ROC-AUC 0.846 [0.838, 0.854]. Then the authors matched on the confounders they had measured. “However, after matching on measured confounders, such as age, gender, and self reported symptoms, our classifiers performance is much weaker (ROC-AUC 0.619 [0.594, 0.644]),” they write. A large part of that apparent skill was being carried by things that were not the respiratory sound, in a cohort of 67,842 people. The only reason anyone knows is that somebody ran the matched comparison.
The uncomfortable part is that measuring the microphones would not have produced either number. Frequency-response deviation and self-noise can be reported in full, for every handset in the study, and every figure can be correct. Those figures describe transducers. They say nothing about what the classifier was tracking.
- The decision a case like this forces is whether microphone response, placement, channel geometry, gain staging, clipping, calibration and metadata get connected all the way through to downstream model behavior, or are left behind as recording details somebody else handled.
- The failure is a confound in the target label, and its magnitude is measurable: 0.846 [0.838, 0.854] unadjusted against 0.619 [0.594, 0.644] once age, gender and self reported symptoms were matched across 67,842 individuals.
- The evidence on file in a study of this kind is the usual pair, frequency-response deviation and self-noise, which describe the transducers and stay clean whatever the classifier has latched onto.
- The practical response is to account for the whole chain rather than the microphone alone: analog filtering, gain, the ADC, compression, denoising and resampling — and to report the matched comparison, not only the headline AUC.
Case
192 apps examined, ten qualified, four tested
Whether a phone can be trusted as an instrument is not a matter of opinion. It is a question with a published answer, and the answer is a funnel. Kardous and Shaw, of NIOSH, tested smartphone sound measurement applications in the Journal of the Acoustical Society of America in April 2014. Their abstract states the shape of it: “A representative sample of smartphones and tablets on various platforms were acquired, more than 130 iOS apps were evaluated but only 10 apps met our selection criteria. Only 4 out of 62 Android apps were tested.”
The NIOSH Science Bulletin gives the same funnel from the other end. Of the 192 apps examined, 10 iOS apps met the criteria for functionality, features and calibration capability. Of those 10, four met the testing criteria, at “± 2 dB mean difference from the reference type 1 sound level meter”. Two of the four came in at mean differences of 0.07 dB unweighted and −0.52 dB A-weighted. The other two were within ±2 dB.
So “four apps were trusted” is the last stage of a long discard. More than 130 iOS candidates were dropped on functional and calibration grounds before any of them met a reference meter, and only 4 of 62 Android apps were tested at all. That is what it costs to establish that a handset is an instrument.
OSHA reaches the same point from the legal side. 29 CFR 1910.95(a) reads: “Protection against the effects of noise exposure shall be provided when the sound levels exceed those shown in Table G-16 when measured on the A scale of a standard sound level meter at slow response.” The column heading in Table G-16 is literally “Sound level dBA slow response”. The table runs from 90 dBA for 8 hours to 115 dBA for ¼ hour or less. The limit is not a bare sound level. The weighting and the time constant are written into the law, which makes it a specification of the instrument as much as of the sound.
Example
Calibrated, not sensitive: the word that bought a decibel
The step from ±2 dB to ±1 dB turned on a single word, and it was bought with hardware. Kardous and Shaw returned to the same four apps in 2016: “The initial study examined 192 apps on the iOS and Android platforms and found four iOS apps with mean differences of ±2 dB of a reference sound level measurement system. This study evaluated the same four apps using external microphones. The results showed measurements within ±1 dB of the reference.” The NIOSH Science Bulletin pins the conditions: “The NIOSH SLM app, when used with an external calibrated microphone, measured sound levels within ± 1 dB of the reference SLM over the testing range of 65 -95 dB SPL in our laboratory”, evaluated for compliance with type 2 requirements of IEC 61672/ANSI S1.4. The decibel came from a microphone bolted on, not from a better setting.
Calibrated, not sensitive, not well specified. Polar pattern, sensitivity, automatic gain control and calibration name four different parts of one capture chain. Two of them can sound alike and still demand different proof, in different units, from different people. So any write-up of that chain has to use these four words exactly.
- Polar pattern is the directional sensitivity of a microphone, which directions it favors and which it attenuates.
- Sensitivity is the electrical output produced for a stated acoustic input, a property of the device on its own — and a handset can be perfectly sensitive while still sitting outside ±2 dB of a reference type 1 sound level meter.
- Automatic gain control is a feedback process that changes gain as signal level changes, which is why a level recorded through it is not a measurement of the sound that arrived.
- Calibration is a traceable mapping from recorded values to a defined reference, which is why it, and not sensitivity, is the word in the NIOSH condition: an external calibrated microphone, over the testing range of 65-95 dB SPL, against a reference SLM.
Comparison
Three normalizations that buy different things
Those three operations were left standing side by side earlier. What separates them is where the answer comes from.
Absolute calibration maps digital values to a physical reference such as sound pressure, and that reference lives outside the file. It is why the NIOSH condition was an external calibrated microphone rather than a better setting on the handset. It is why OSHA had to name the A scale and slow response in the regulation itself.
Relative normalization stays inside the recording. That does not make it vague; it is standardised to the decimal place. ITU-R BS.1770 defines the measurement in LKFS/LUFS. EBU Recommendation R 128, first published in February 2010 and revised in November 2023, recommends “that the Programme Loudness Level shall be normalised to a Target Level of −23.0 LUFS. Where attaining the Target Level is not achievable practically (for example, live programmes), a tolerance of ±1.0 LU is permitted.” Quality-control workflows are held to ±0.2 LU, with a true-peak ceiling of −1 dBTP. A target, a tolerance and a ceiling — and still a number referenced to full scale rather than to sound pressure.
Domain normalization stays inside as well, and is doing something else again. It reduces device variation for a model, at the risk of removing part of the task signal with it. Each of the three has to be shown separately. An operation that never leaves the recording, ±0.2 LU included, cannot tell anyone that the recorder was the thing being heard.
Absolute calibration
Maps digital values to a physical reference such as sound pressure.
- Decision focus: Characterize the transducer
- Useful evidence: Frequency-response deviation and self-noise
- Watch for: Device identity leaking into the target label
- Best used when its assumptions are documented for microphones, channels, gain, clipping, and calibration
Relative normalization
Rescales recordings to a common statistic without restoring physical units.
- Decision focus: Specify placement and geometry
- Useful evidence: Clipping, saturation, and automatic-gain incidence
- Watch for: Automatic gain control changing temporal dynamics
- Best used when its assumptions are documented for microphones, channels, gain, clipping, and calibration
Domain normalization
Attempts to reduce device variation for a model, with possible loss of task signal.
- Decision focus: Set gain with headroom
- Useful evidence: Channel delay and geometry error
- Watch for: Channel swaps or unsynchronized array recordings
- Best used when its assumptions are documented for microphones, channels, gain, clipping, and calibration
Visual
Keep the firmware version with the file
The workflow runs in one direction. Characterize the transducer — measure frequency response, self-noise, directivity, sensitivity and nonlinear range. Specify placement and geometry — record distance, orientation, mounting, channel order and synchronization. Set gain with headroom — test quiet speech and worst-case transients while monitoring automatic gain behavior. Preserve calibration metadata — store device, firmware, channel, unit and calibration history with each recording.
What you measured about the transducer does not stay in the measurement report. The assumptions baked into how you characterized it are consumed downstream, when you specify placement and geometry, and nothing later reopens them. Every step after the first one trusts the first one.
That leaves a single place where a mistake can still surface. The calibration metadata kept with the file, device and firmware version included, is the part of the characterization that reaches whoever asks the awkward question months later. Discard it after ingestion — the fourth failure mode on the opening list — and there is nothing left to check the first step against. Preserve it, and a repair remains possible. CHiME-6 could compensate for frame-drops and clock-skew across 32 microphones per session two years after the recordings were made. It could do that because the audio and its device structure had survived intact.
Characterize the transducer
Measure frequency response, self-noise, directivity, sensitivity, and nonlinear range.
Specify placement and geometry
Record distance, orientation, mounting, channel order, and synchronization.
Set gain with headroom
Test quiet speech and worst-case transients while monitoring automatic gain behavior.
Preserve calibration metadata
Store device, firmware, channel, unit, and calibration history with each recording.
Whatever you measured about the transducer is trusted by every later step, and the calibration metadata is the only place a mistake can still surface.
Analogy
A camera lens with an invisible color cast
If a lens tinted each class a different color, image recognition would learn to identify the lens rather than the object. The training curves would look excellent while it did. That is the shape of a 0.846 that becomes 0.619 the moment the measured confounders are matched. The first number was never a lie. It was an answer to a different question.
The comparison flatters the microphone, though. A tint changes color and nothing else. A microphone also alters timing, directionality, gain and the reverberant mixture. That gives the capture hardware four more ways to become the feature, and a team four more things to check before concluding that the shortcut is gone.
Capture hardware is part of the dataset and must be governed as such.
Example
72.8% on device A, 54.6% on device S2
Robustness across devices was written into the task from the start. “The systems are expected to be robust to different devices”, says the DCASE 2020 Challenge Task 1 description. Then the challenge published the table that shows what robust costs.
Task 1A used TAU Urban Acoustic Scenes 2020 Mobile: 64 hours of audio, 23,040 ten-second segments, 10 acoustic scenes, 10 cities and 9 devices — real A, B and C plus simulated S1–S6. The full dataset spans 12 European cities, two of which appear only in the 11-device evaluation set. The device-wise results table for the baseline reads 72.8% on device A, 68.9% on C, 62.7% on S1, 61.7% on B, 58.2% on S3 and 54.6% on S2. Same model, same scenes. That is an 18.2-point spread across the devices seen in training. On device D, which the model had not seen, it was 22.8%. The baseline's overall evaluation accuracy was 51.4% (50.5–52.3), against 76.5% (75.8–77.3) for the top system, Suh_ETRI_task1a_3.
One mean accuracy reports 51.4% and hides every one of those numbers. Frequency-response deviation and self-noise, meanwhile, would have described nine well-made capture paths. They would have said nothing about the 18.2 points.
- Start with the transducers themselves, which means frequency-response deviation and self-noise for each device in the study — necessary, and on its own carrying no release claim.
- Then how the chain behaved while recording, which is clipping, saturation and automatic-gain incidence.
- Then the hard slices, built from channel delay and geometry error and from recordings where the device rather than the sound predicts the label — the CHiME frame-dropping and clock-drift case is what an unmeasured version of this looks like.
- Then the lifecycle question, which is performance sliced by device, by where the microphone sat, by room and by firmware, reported the way DCASE reported it: 72.8, 68.9, 62.7, 61.7, 58.2 and 54.6 on the devices seen in training and 22.8 on one that was not, rather than a single 51.4%.
A device-wise table turns "robust to different devices" into 72.8% against 54.6%, and 22.8% on a device the model had never heard.
Steps
Create a capture-chain datasheet
The format already has a published proposal behind it, and the exercise is easier if you use it. In 2018 Gebru and six co-authors proposed that every dataset travel with a datasheet, and argued it by analogy: “In the electronics industry, every component, no matter how simple or complex, is accompanied with a datasheet that describes its operating characteristics, test results, recommended uses, and other information. By analogy, we propose that every dataset be accompanied with a datasheet that documents its motivation, composition, collection process, recommended uses, and so on.” Motivation, composition, collection process, recommended uses: four headings, and a capture chain fits all four.
Write the steps out in order. List every transformation, including analog filtering, gain, the ADC, compression, denoising and resampling. Map confounders — which devices, rooms and placements correlate with labels. That is the question whose unmatched answer was 0.846 and whose matched answer was 0.619. Design counterbalanced collection, spreading target classes across devices and environments. Then define rejection rules.
The page is worth writing only if a stranger can use it. It earns its place when another team can pick it up and ask, without your help, whether the recorder rather than the sound is doing the work in your results. And the part that gets skipped is the last one. The page has to name the point at which a recording is too clipped, too unsynchronized, or too undocumented to keep.
1. List every transformation
Include analog filtering, gain, ADC, compression, denoising, and resampling.
2. Map confounders
Ask which devices, rooms, and placements correlate with labels.
3. Design counterbalanced collection
Spread target classes across devices and environments.
4. Define rejection rules
State when a recording is too clipped, unsynchronized, or undocumented to use.
Until it names a threshold for throwing a recording away, the page is a description of your equipment rather than a decision anyone can be held to.
Key takeaways
- Normalization aligns average scale and nothing more. Saturation, blocked frequencies, device response and a source masked by reverberation and competing sound all happened before any scaling factor existed — and the channel misalignment CHiME-6 had to repair it never touches at all.
- Confounding in audio has a measured size. Across 67,842 individuals, 23,514 of them PCR-positive, COVID-19 audio classifiers scored ROC-AUC 0.846 [0.838, 0.854] unadjusted. After matching on age, gender and self reported symptoms it was 0.619 [0.594, 0.644], reported in Nature Machine Intelligence in February 2024.
- Treat a phone as a measuring instrument only where that has been demonstrated. Kardous and Shaw examined 192 apps, found 10 iOS apps meeting the functionality, features and calibration criteria, and tested 4 within “± 2 dB mean difference from the reference type 1 sound level meter”. The 2016 follow-up reached ±1 dB only by using external calibrated microphones, over the 65-95 dB SPL testing range.
- Characterize the microphone and its placement up front, because placement and geometry decisions consume that characterization and never reopen it. CHiME-6 could compensate for audio frame-dropping and clock-drift across 32 microphones per session two years later only because the recordings and their device structure survived.
- Absolute calibration answers to a physical reference outside the file, as OSHA does when 29 CFR 1910.95 heads its Table G-16 column “Sound level dBA slow response”. Relative normalization can be exact and still never leave the recording, as EBU R 128 is at −23.0 LUFS with ±1.0 LU and a −1 dBTP ceiling. Domain normalization is a third thing again, and none substitutes for another.
- Frequency-response deviation and self-noise describe the transducers. Only performance sliced by device says whether the transducer was what the model heard — the DCASE 2020 Task 1A baseline ran 72.8% on device A and 54.6% on S2, 22.8% on unseen device D, behind one overall figure of 51.4%.