Skip to content
AI.info

Speech and audio

Sound Waves, Amplitude, Phase, and Decibels

Build a practical physical model of pressure waves, frequency, amplitude, phase, power, and logarithmic level measurements.

By the end you can

A decibel without its reference means nothing

Sound is a variation in time, and three quantities describe that variation. Frequency is the rate at which a pattern repeats. Amplitude is the size of the variation, always a size relative to some reference. Phase is where in its cycle a component sits. The three interact. No single scalar stands in for all of them, perceptually or operationally.

A decibel is not a fourth quantity. It is a ratio between something measured and something chosen as a reference, written on a logarithmic scale. That makes any dB figure half a statement. The other half usually goes unwritten: which physical quantity was measured, against which reference, with what weighting, over what bandwidth, averaged across what window.

The missing half can be put as a number. Sound pressure in air and other gases is referenced to 20 µPa. Sound pressure in water and other liquids is referenced to 1 µPa. ISO 1683:2015 fixes both, and notes that the 1 µPa reference gives a level approximately 26,0 dB greater for the identical sound pressure. Nothing about the pressure changed. The number moved because the denominator did. The standard states the obligation that follows: “The reference value used to establish a level for a certain acoustical quantity should always be stated together with the respective level.”

Digital full scale, sound-pressure level and power ratios all arrive spelled the same way, and none of them converts into another. Whoever builds this has one thread to hold from end to end. The physical quantity is named at the moment of measurement. It has to still be the quantity in hand when the number is finally read in context, with somebody answerable at each end.

The same sound pressure reads about 26,0 dB apart depending on whether the reference is 20 µPa in air or 1 µPa in water, and a bare dB figure does not say which. ISO 1683:2015 fixes both.

Case

Two agencies, one workplace, two noise limits

The missing half is not a hypothetical. Two United States federal agencies write it differently for the same workplace. Neither of them is doing anything wrong.

OSHA sets the limit in 29 CFR 1910.95, and the rule specifies its own measurement: “Protection against the effects of noise exposure shall be provided when the sound levels exceed those shown in Table G-16 when measured on the A scale of a standard sound level meter at slow response.” The A scale, a standard sound level meter, slow response — all three are in the rule, not left to whoever holds the meter. Table G-16 then allows 8 hours at 90 dBA, 6 h at 92, 4 h at 95, 3 h at 97, 2 h at 100, 1.5 h at 102, 1 h at 105, 0.5 h at 110 and 0.25 h or less at 115 dBA. The permitted time halves for every 5 dBA. Two further clauses in the same standard carry different numbers for different purposes. One sets an 85 dBA 8-hour time-weighted-average action level, a 50 percent dose, which triggers a hearing conservation program rather than a violation. The other caps impulsive or impact noise at 140 dB peak sound pressure level. One agency, one regulation, three numbers, none of them interchangeable.

A different agency reaches a different answer for the same room. NIOSH recommends 85 dBA as an 8-hour time-weighted average, published in 1998. The five-decibel gap against OSHA's 90 dBA is the visible disagreement, and it is the smaller of the two. The larger one is the exchange rate: how fast the permitted time shrinks as the level rises. NIOSH states it in its own words: “For each 3 dBA increase in noise level, NIOSH recommends reducing the exposure duration by half.” Table G-16 halves for every 5 dBA instead, which is why OSHA's rows run 8 hours at 90, 4 h at 95, 2 h at 100.

Same sound, same A-weighting, same eight hours. Two thresholds, and two rules for what to do about them. The exchange rate is part of the measurement, not a footnote attached to it. A room does not get louder when you change which agency you are reading. What changes is what the figure 90 was counting all along.

Figure

Two agencies, one workplace, two answers: what the OSHA and NIOSH halving rules allow at 90, 95 and 100 dBA once each is applied.

Example

Same algorithm, two loudness targets, different material measured

OSHA and NIOSH at least print their averaging rules beside their numbers, and they are separate agencies with separate mandates. The harder version of the problem appears when two bodies share the measurement algorithm and still cannot exchange their numbers.

Both broadcast loudness regimes measure with ITU-R BS.1770. EBU R 128-2023 normalises Programme Loudness to −23.0 LUFS. It permits ±1.0 LU for live programmes and ±0.2 LU in quality control, and sets a True Peak ceiling of −1 dBTP. ATSC A/85:2013 sets a Target Loudness of –24 LKFS, with ±2 dB anticipated measurement variation and a true-peak level below −2 dBTP.

The targets are 1 LU apart, which looks like a rounding difference. That is not the interesting part. The interesting part is what each number was computed over. EBU R 128 normalises the programme. A/85 measures the Anchor Element, usually dialog, rather than the whole programme. Two files can therefore be compliant with their respective rules, described in near-identical units, and still not comparable. One figure summarises everything that was broadcast. The other summarises the voice inside it.

A window, and the material it is drawn over, is not a formatting detail. It is a decision about what is allowed to disappear before anyone sees the result.

  • What is at stake across this lesson is a practical physical model of pressure waves. It has to hold frequency, amplitude, phase, power and logarithmic level together, instead of consulting them one at a time.
  • The failure that starts most of the others is adding decibel values directly, without checking what quantity each of them measured, against which reference, over which material.
  • What separates −23.0 LUFS from –24 LKFS is not a tighter number but a declared procedure: peak, RMS and energy over windows somebody wrote down in advance, the way ITU-R BS.1770-5 does.
  • The cheapest habit that prevents the whole sequence is asking what zero decibels means in the measurement you are holding — 20 µPa, 1 µPa, digital full scale, or something else again.

Example

Who needs which quantity

Whether a difference like that matters depends entirely on who is reading the number. Sometimes the answer is fixed by law rather than by taste. The CALM Act made the ATSC A/85 Recommended Practice binding on United States television broadcasters: “Effective December 13, 2012, television broadcast stations must comply with the ATSC A/85 RP incorporated by reference, see § 73.8000), insofar as it concerns the transmission of commercial advertisements.” A measurement algorithm, a reporting unit and a choice of which material to measure were written into federal rules, with a date on them.

Outside the law the same quantities get read toward different ends. Hearing and acoustics, music production, predictive maintenance and spatial capture all use them. Each of the four needs a different one to survive the summary.

  • Hearing and acoustics increasingly does not track a level at all. It tracks an energy budget. The WHO–ITU global standard for safe listening devices, ITU-T H.870, sets Mode 1 for adults at 1.6 Pa²h per 7 days, derived from 80 dBA for 40 hours a week. Mode 2, for sensitive users such as children, is 0.51 Pa²h per 7 days, derived from 75 dB for 40 hours a week. It is an equal-energy rule, in the standard's own words: “listening to a 100 dB sound for 16 minutes will have the same impact as listening to an 80 dB sound for 40 hours”.
  • Music production and delivery read peak, RMS, crest factor and phase toward different decisions, and the delivery target is written down rather than felt: “For delivery or exchange of content without metadata (and where there is no prior arrangement by the parties regarding loudness), the Target Loudness value should be –24 LKFS.” That figure says nothing on its own about the −1 dBTP or −2 dBTP ceiling sitting beside it.
  • Predictive maintenance is where a short impact has to survive the summary. Harmonics and transients can reveal different failure mechanisms. A gated or long-window average is designed to discard exactly the quiet-then-loud structure that carries one of them.
  • Spatial capture lives on inter-channel phase and delay. They carry the localization information, and they are the first thing a level summary throws away — as BS.1770's channel-weighted summation does when it collapses channels into one figure.

Key idea

Decibels do not add

OSHA's Table G-16 and NIOSH's recommended exposure limit share a unit, a weighting and an eight-hour averaging period. They are still not the same measurement: 90 dBA against 85 dBA, halving every 5 dBA against halving every 3 dBA. EBU R 128-2023 and ATSC A/85:2013 share an algorithm, ITU-R BS.1770, and are still not the same measurement: −23.0 LUFS over the programme against –24 LKFS over the Anchor Element. Two decibel figures drawn from different references or different windows are further apart than either pair, and nothing in their appearance says so. That is how a healthy-looking level summary hides this failure completely. The figures behind it were added together without anyone checking what each of them measured.

The other three failures are the same move in different clothes. Peak amplitude gets read as perceived loudness. It is a statement about the single largest instant, not about how the whole thing is heard. That is why A/85 carries a separate true-peak limit below −2 dBTP alongside its –24 LKFS target instead of deriving one from the other. Phase gets discarded at exactly the point where transients or spatial cues are what the question turns on. And a calibrated acoustic level gets set beside uncalibrated digital sample values, two scales counting from different zeros — the way 20 µPa and 1 µPa count from different zeros and put about 26,0 dB between two readings of one pressure.

None of the four leaves a mark on the output. A summary looks just as healthy built from four incompatible references as from one. So the check belongs before the figures are combined, not after the total starts looking strange.

Two decibel figures can be compared only once you know what each measured, against which reference and over which window. Without that, adding them yields a number describing nothing.

Visual

Four choices before a level means anything

Four separate choices stand behind any stated audio level: measurement, modeling, decision and verification. The path below keeps them apart instead of hiding them inside one score.

Measurement is where the physical quantity, the reference and the weighting get fixed. That means 20 µPa or 1 µPa under ISO 1683:2015, the A scale on a standard sound level meter at slow response under OSHA's rule, K-frequency weighting under ITU-R BS.1770-5.

Modeling is where the averaging gets fixed: the window, the exchange rate, and whether phase survives at all. BS.1770-5 uses 400 ms blocks overlapping by 75%. The two workplace rules use an eight-hour time-weighted average, with a 5 dBA or a 3 dBA exchange rate. A/85 uses the Anchor Element rather than the whole programme.

Decision is where the resulting number is set against a limit — Table G-16, 140 dB peak sound pressure level, −1 dBTP, 1.6 Pa²h per 7 days.

Verification is where somebody asks whether the first two choices suited the question being asked. It is the step most often skipped. By then the number has stopped looking like a choice and started looking like a fact.

FigureTimeline · 4 stops
  1. 1. Identify the physical quantity

    Determine whether the sensor measures pressure, voltage, acceleration, or a derived digital value.

  2. 2. Select a reference

    State the reference level used by the ratio or calibration.

  3. 3. Choose time and frequency resolution

    Decide whether a transient, tone, or long average matters.

  4. 4. Interpret in context

    Relate the measurement to audibility, clipping, equipment health, or another defined outcome.

A number read later carries the physical quantity and the reference you settled on at the start, and no step in the path stops to ask whether that choice was right.

Example

Four quantities, none of them loudness

Each of these four has already done work in this lesson, and not one of them is another name for loudness. Each may ask for its own proof, arrive in its own unit and put the decision in different hands. So anything written about them has to use the four names exactly.

  • Frequency is the rate of repetition, measured in cycles per second or hertz. It is a count of cycles, and on its own no statement at all about how loud anything is. It does decide whether A-weighting, or the K-frequency weighting of BS.1770-5, raises or lowers the figure you end up quoting.
  • Phase is the relative position of a periodic component within its cycle. It is precisely what a channel-weighted, gated level summary discards on its way to a single number.
  • RMS is a root-mean-square summary related to signal power under stated assumptions. Integrating that power over time gives an energy rather than a level. ITU-T H.870 budgets exposure in pascal-squared-hours, 1.6 Pa²h per 7 days for Mode 1 and 0.51 Pa²h per 7 days for Mode 2. No meter reading in dBA can be substituted for that unit.
  • dBFS is a digital level referenced to the maximum representable full-scale value, which is why it cannot be laid beside a 90 dBA row of OSHA's Table G-16. The two count from different zeros.

Example

Name the window, or the number is empty

Naming those four quantities is not yet reporting them. Peak, RMS and energy over declared windows are the figures worth reporting. On their own they still cannot tell you whether the decibels behind them were ever comparable.

A declared window, written down properly, looks like this. ITU-R BS.1770-5 fixes K-frequency weighting, a per-channel mean square, and a channel-weighted summation that excludes the LFE channel. Then comes “gating of 400 ms blocks (overlapping by 75%), where two thresholds are used:” — an absolute threshold at −70 LKFS, and a relative one at −10 dB below the level measured after the first has been applied. Every one of those choices decides what is allowed to disappear. The gates stop silence and near-silence from dragging the figure down, which also means the specification says in advance which material will not be counted. That is the standard to hold your own reporting to. Not the exact numbers. The fact that the numbers exist in writing, and anyone can recompute against them.

So each figure travels with the rest of its statement. Say which unit it is in, over which recordings, how uncertain it is, and what the equipment was doing at the time. Put frequency and phase response beside it where those are relevant.

  • The core evidence is peak, RMS and energy over declared windows. Declaring the window the way BS.1770-5 does — 400 ms blocks, 75% overlap, thresholds at −70 LKFS and at −10 dB below the first-pass level — is what makes your three figures comparable with anybody else's three.
  • System behaviour needs frequency and phase response wherever they bear on the question. A level on its own says nothing about either, and a weighting curve applied before the average has already reshaped the level you are quoting.
  • The robustness slice is a calibrated level carrying its reference and its weighting in the same breath as the number. That is the whole difference between 90 dBA measured on the A scale at slow response and a bare 90.
  • Over the life of a transient-rich signal, dynamic range and crest factor are the figures in which a short impact shows up at all. It is the thing a gate at −70 LKFS, or a peak cap at 140 dB peak sound pressure level, treats as a separate question from the average.

A level quoted without its uncertainty and without the state of the equipment describes one afternoon in one room, not the machine.

Steps

Audit a level claim

Take a level claim, yours or somebody else's, and put the questions to it in order. What physical quantity was measured? Against which reference — 20 µPa, 1 µPa, digital full scale? With what weighting and bandwidth, A-weighted or K-weighted? Over what window, and over which material: everything, or an Anchor Element inside it? And what exchange rate, if any, turns that level into a permitted time? That last one is where OSHA's 5 dBA and NIOSH's 3 dBA diverged while appearing to agree on a unit.

Then turn the same questions on your own work. Note what your choice of physical quantity assumes, one case where that assumption fails, and what changes once the number is read in context. ISO 1683:2015 already tells you the minimum: state the reference value together with the level. A level claim survives the audit only when another team has tried to break it by asking what quantity each decibel figure actually measures.

FigureProcess · 4 steps
  1. 1. Find the reference

    Ask what zero decibels means in the measurement.

  2. 2. Find the window

    Identify peak, short-term, or long-term aggregation.

  3. 3. Find the weighting

    Record any frequency or temporal weighting.

  4. 4. Test a counterexample

    Construct two signals with similar level but different operational meaning.

Pair two signals that read the same on a meter and behave differently in the room, and the number stops being an argument on its own.

Key takeaways