Speech and audio
Pitch, Loudness, Timbre, and Auditory Perception
Explain pitch, fundamental frequency, loudness, timbre, masking, critical bands, and the limits of perceptual proxies.
By the end you can
- Define pitch, loudness, timbre, and auditory perception as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish fundamental frequency, loudness measure, and timbre representation without treating them as interchangeable
- Trace the workflow from name the perceptual question through triangulate results
- Evaluate pitch, loudness, timbre, and auditory perception using pitch error and voicing accuracy for controlled tasks and evidence from difficult deployment slices
Example
The integrated level matched and the balance broke
Normalising every clip to a single loudness target is not a hypothetical pipeline decision. In the United States it is law. The CALM Act, approved 15 December 2010, gave the FCC one year to make the ATSC A/85 Recommended Practice mandatory for commercial advertisements. The rule that followed carries its own date: "Effective December 13, 2012, television broadcast stations must comply with the ATSC A/85 RP incorporated by reference, see § 73.8000), insofar as it concerns the transmission of commercial advertisements." That is 47 CFR § 73.682(e)(1). The unbalanced parenthesis is in the original.
So the target is real, dated and enforceable. A pipeline that matches it has done what the regulation asks. On paper it works: the integrated level comes out the same for every clip. In the room it does not follow. Quiet ambience is raised until masking and background noise dominate it. A transient-rich performance loses headroom while its integrated level matches everything else.
Matching the integrated level is not the same as matching what a listener hears. A legal target does not measure the second thing. A review that asked only for the usual pitch and voicing numbers on controlled tasks would have recorded none of it. Those numbers are real. They answer a question nobody was having trouble with.
- A case like this puts a whole vocabulary at stake: pitch and fundamental frequency, loudness, timbre, masking, critical bands. It also puts a limit on that vocabulary — the point where a perceptual proxy stops standing in for what somebody actually heard.
- One failure runs narrower than the vocabulary, and this lesson keeps returning to it: inferring emotion or personality from pitch alone. The EU AI Act prohibits the practice outright in workplaces and education institutions.
- The evidence usually offered in these arguments is pitch error and voicing accuracy on controlled tasks. CREPE reports 0.999 +/- 0.002 on one dataset and 0.909 +/- 0.126 on another.
- When someone hands over a vague quality score, ask for a concrete comparison instead: which dataset, which threshold, which listeners, which tolerance.
Analogy
Same brightness, different color
The trap has a visual version that is easier to see. Two paints can measure the same brightness and still look nothing alike, because hue, saturation, and the colors around them differ. Brightness is a real measurement. It simply does not determine the object.
Acoustic level alone cannot determine the auditory object either. An identical integrated level across a catalogue guarantees nothing about how those clips sit beside one another — even when that level has been mandated since 13 December 2012.
The analogy stops in one place, and the place it stops is the interesting one. Paint sits still while it is being looked at. Sound unfolds over time, and hearing adapts as it goes.
Perceptual labels require controlled human evidence, not only signal statistics.
A perception frozen into a specification
The quantities that matter here are perceptual ones. Pitch, loudness, and timbre are related to frequency, level, and spectrum without being identical with them. Human hearing integrates context, duration, masking, what the listener expects, who the listener is, and how the sound is played back. A psychoacoustic model predicts how a named group of listeners behaves under named conditions. It does not establish that every listener hears the same object.
That is not a critic's complaint about the standards. It is printed inside them. ITU-R BS.1770-5 was approved in November 2023. It specifies the K-weighting/mean-square/channel-weighted/gated algorithm that turns audio into a loudness number in LKFS, gated at -70 LKFS and at -10 dB relative to the first-pass level. Having defined all of that, it adds: "NOTE 1 – Users should be aware that measured loudness is an estimation of subjective loudness and involves some degree of uncertainty depending on listeners, audio material and listening conditions." That is page 2. Before the algorithm is even finished, the standards body has told you that its own number depends on who is listening and to what.
The freezing of a perception into a specification can be watched happening. The equal-loudness contour is the model that maps physical level onto perceived loudness, and it is ISO 226. That standard was fully revised in 2003, after errors reported in 1985. It was revised again as ISO 226:2023, twenty years later. A 2024 account by Suzuki and colleagues gives the reasons: "One motivation for the revision was to reflect the lowering of the threshold of hearing at 20 Hz by 0.4 dB in ISO 389-7:2019. In addition, the following two points of substance were revised: (1) implementation of the power exponent relating loudness perception to physical intensity formulated in an academic paper published in 2004, which describes the derivation of ELLCs relating to ISO 226:2003, and (2) adoption of mathematical expressions that preserve the appropriate number of significant digits." The 2023 contours differ from the 2003 ones by at most 0.6 dB. Small — and the size of a disagreement that took twenty years and an international committee to settle.
That is where a single loudness target comes from. It is also why a pipeline has every reason to trust the one it uses. What comes with such a number, unwritten, is its scope: a listening judgment made by some group under some conditions. Applied afterwards to ambience, to a transient-rich performance, and to whoever happens to be playing them back.
ITU-R BS.1770-5 prints the caveat on page 2 — measured loudness is an estimation that depends on listeners, material and conditions — and teams still stop asking whose ears it came from.
Example
Pitch is heard; F0 is measured
The standards survive that scope problem by being exact about what they name. EBU R 128 was first published in 2010 and is now in its November 2023 version. It does not say "make it loud enough". It recommends "that the Programme Loudness Level shall be normalised to a Target Level of −23.0 LUFS. Where attaining the Target Level is not achievable practically (for example, live programmes), a tolerance of ±1.0 LU is permitted." Quality-control workflows get ±0.2 LU. The maximum True Peak Level is -1 dBTP, with a ±0.3 dB measurement tolerance. And the signal must be measured in its entirety, "without emphasis on specific foreground elements". Named quantity, named unit, named target, named tolerance, named measurement scope.
Reports about pitch are rarely so careful. Pitch, masking, and timbre are things a listener experiences. Fundamental frequency is the measurement people reach for instead. Write "pitch" where a system measured fundamental frequency and three things move at once: what now has to be shown, the unit it has to be shown in, and who is entitled to rule on it. Keep the perceived quantity and the measured one apart. Name each as exactly as R 128 names -23.0 LUFS.
- Fundamental frequency, often written F0, is the repetition rate associated with a periodic source. It is measured — and measured with an accuracy that depends on the material, as CREPE's 0.999 against 0.967 shows.
- Pitch is the perceptual ordering of sounds from low to high — where a listener puts a sound, not what an estimator returns.
- Masking is the reduced audibility of one sound because of another. It is what swallowed the ambience once the pipeline raised everything to a single target level.
- Timbre covers the perceived qualities that distinguish sounds beyond pitch, loudness, and duration. It is the part of the balance normalisation breaks while the -23.0 LUFS figure holds steady.
Key idea
Pitch does not carry personality
Four habits push these measures past what they can hold, and the case above is one of them.
Start with the one this lesson keeps naming: inferring emotion or personality from pitch alone. This is no longer only a methodological objection. Under Article 5(1)(f) of the EU AI Act, Regulation (EU) 2024/1689, inferring emotions from a person in workplaces and education institutions is a prohibited practice. It has applied since 2 February 2025, under Article 113(a). Recital 44 gives the legislature's stated grounds: "There are serious concerns about the scientific basis of AI systems aiming to identify or infer emotions, particularly as expression of emotions vary considerably across cultures and situations, and even within a single individual. Among the key shortcomings of such systems are the limited reliability, the lack of specificity and the limited generalisability." Culture, situation, and variation within one person. Three ways the acoustic signal fails to fix the mental state.
Second, using integrated loudness to judge brief safety-critical alerts. R 128 asks for the signal to be measured in its entirety, "without emphasis on specific foreground elements". That is exactly the wrong instrument for a sound that is over in a moment. Third, ignoring hearing ability, language, culture, and playback device. Fourth, optimizing a perceptual proxy without validating the actual user task.
All four stretch a model fitted to one named group of listeners under named conditions across an audience nobody has named: the hard-of-hearing user, the second-language listener, the phone speaker held at arm's length. Stretched that far, the score stops describing hearing at all. It cannot say which rendering people prefer, whether the alert was understood, whether the speaker was distressed, or whether a clinician should act.
Pitch tells you about a signal, not about a person — and since 2 February 2025, Article 5(1)(f) of the AI Act treats reading emotion off a person at work or in school as prohibited, not merely unsupported.
Comparison
Periodicity, level, and everything else
Behind all four habits sits one substitution: treating three quantities as though any of them could stand in for the others.
Fundamental frequency is a physical periodicity estimate that may correlate with perceived pitch. It is the thing CREPE reports to three decimal places, with a standard deviation attached and a dataset named beside it. A loudness measure answers a different question. It comes in LKFS or LUFS, gated at -70 LKFS and at -10 dB below the first pass, and reported against a target such as -23.0 LUFS. A timbre representation answers a third question, and it is the one with no standard number at all.
They rest on different things entirely. That is how a pipeline can hold one of them fixed to within ±1.0 LU and lose another without a single number moving.
Fundamental frequency
A physical periodicity estimate that may correlate with perceived pitch.
- Decision focus: Name the perceptual question
- Useful evidence: Pitch error and voicing accuracy for controlled tasks
- Watch for: Inferring emotion or personality from pitch alone
- Best used when its assumptions are documented for pitch, loudness, timbre, and auditory perception
Loudness measure
A model of perceived level under defined weighting and integration.
- Decision focus: Choose acoustic proxies
- Useful evidence: Loudness and level measures with reference and integration stated
- Watch for: Using integrated loudness for brief safety-critical alerts
- Best used when its assumptions are documented for pitch, loudness, timbre, and auditory perception
Timbre representation
A description of spectral and temporal qualities beyond pitch and level.
- Decision focus: Design listening evidence
- Useful evidence: Listening-test reliability and confidence intervals
- Watch for: Ignoring hearing ability, language, culture, and playback device
- Best used when its assumptions are documented for pitch, loudness, timbre, and auditory perception
Steps
Design a perceptual listening test
The person who designed the test catches none of this. Hand it to a colleague who did not, and see whether they can use it to challenge the claim that pitch alone reveals emotion or personality. If they cannot, the test is not evidence. It is a formality.
The counterexample that exercise needs already exists, with a number on it. Human listeners judging emotion in a voice are right about 70% of the time when they are handed five labels to choose between. The figure comes from a 2003 meta-analysis by Juslin and Laukka in Psychological Bulletin, pooling 39 studies of vocal expression and 12 of music performance across 73 decoding experiments: "Overall decoding accuracy across within-cultural vocal expression and music performance was .89, which is equivalent to a raw accuracy score of .70 in a forced-choice task with five response alternatives". Across cultures it falls: pi = .84 against .90 within-culture. That is the honest ceiling for the best case — human listeners, forced choice, a short closed label set. Any system reading emotion off pitch alone is claiming to beat it.
Three things belong in the record: what your statement of the perceptual question assumes, one counterexample to it, and what a set of triangulated results would oblige you to change. The last is what makes the exercise binding. It names the outcome that would change your mind before you have seen it.
1. Define one judgment
Ask for a concrete comparison rather than a vague quality score.
2. Control presentation
Set playback device, level, order, randomization, and environment.
3. Recruit the right listeners
Match language, expertise, hearing profile, and user population.
4. Analyze uncertainty
Report listener and item variation, not only a grand mean.
If trained human listeners decode vocal emotion at about .70 in a five-way forced choice, and lower across cultures, a grand mean from your own listening test is hiding the disagreement that matters.
Example
Four numbers, not one
The pipeline shipped on one number. A release claim needs four, reported together.
Start with the pitch and voicing numbers from the controlled tasks, and see how little a single one of them settles. CREPE's raw pitch accuracy is 0.999 +/- 0.002 on the 6.16-hour RWC-synth set. On the MDB-stem-synth set — 230 tracks, 15.56 hours, 25 instruments — it is 0.967 +/- 0.091. Both figures are at a 50-cent threshold. Tighten the threshold to 10 cents and the second falls to 0.909 +/- 0.126. The authors named the cause themselves in 2018: "On the RWC-synth dataset, CREPE yields a close-to-perfect performance where the error rate is lower than the baselines by more than an order of magnitude. While these high accuracy numbers are encouraging, those are achievable thanks to the highly homogeneous timbre of the dataset." Timbral diversity is what separates 99.9% from 90.9%. Timbre is the quantity the headline number does not mention.
Then report the loudness and level measures with their reference and integration stated. Those are the two things a single target leaves implicit, and the two the pipeline never had to declare. Report the hard slices with their listening-test reliability and confidence intervals — the slices where someone might be tempted to draw a personality claim out of pitch alone. Then report the task-specific result: intelligibility, preference, or detection.
The pitch and voicing score carries a release claim only when the other three are reported beside it.
- The core task evidence is pitch error and voicing accuracy on controlled tasks. Each figure carries its dataset, its cent threshold and its variance — 0.999 +/- 0.002 and 0.909 +/- 0.126 are the same system.
- The system behavior is the loudness and level measures, with their reference and integration stated: LKFS or LUFS, the -70 LKFS and -10 dB gates, the target and the tolerance.
- The robustness slice is listening-test reliability, reported with its confidence intervals, against a human benchmark that is itself about .70 in a five-way forced choice.
- The lifecycle evidence is task-specific intelligibility, preference, or detection performance.
Report pitch error and voicing accuracy for controlled tasks together with task-specific intelligibility, preference, or detection performance — a tracker that scores 0.999 on homogeneous timbre and 0.909 on diverse timbre has no single accuracy to quote.
Key takeaways
- Pitch, loudness and timbre are perceptual constructs, related to frequency, level and spectrum but never identical with them — which is why ITU-R BS.1770-5 warns on page 2 that measured loudness is an estimation of subjective loudness.
- A psychoacoustic model estimates behavior for specified populations under specified conditions, and those specifications get renegotiated. ISO 226 was revised in 2003 and again in 2023, partly for a 0.4 dB change to the 20 Hz threshold in ISO 389-7:2019.
- Name the perceptual question first and triangulate the results at the end. Otherwise a legally mandated target — 47 CFR 73.682(e)(1), effective 13 December 2012 — will quietly stand in for what a listener actually hears.
- Fundamental frequency, a loudness measure and a timbre representation answer related but different questions. -23.0 LUFS ±1.0 LU is exact about loudness and says nothing about the timbral balance that CREPE's 0.999-to-0.909 gap is made of.
- Reading emotion or personality off pitch alone treats one acoustic measure as a person. Article 5(1)(f) of the EU AI Act prohibits it at work and in education from 2 February 2025, on grounds of limited reliability, lack of specificity and limited generalisability.
- Report pitch error and voicing accuracy for controlled tasks together with task-specific intelligibility, preference or detection performance. One number is what the loudness pipeline had, and Juslin and Laukka's .70 in a five-way forced choice shows how little one number settles.