Skip to content
AI.info

Speech and audio

Spoken Language, Accent, and Paralinguistic Analysis

Cover language and accent recognition, emotion and affect labels, age and state proxies, annotation ambiguity, fairness, and appropriate boundaries.

By the end you can

Example

A Seasonal Manager promotion, an AI-scored video interview, and a complaint filed on 19 March 2025

A Deaf Pawnee woman was denied a promotion to Seasonal Manager after an AI-scored video interview. On 19 March 2025 the ACLU and three co-counsel filed a complaint about it. It went to the Colorado Civil Rights Division and the EEOC, against Intuit and HireVue. It is brought on behalf of the applicant, D.K. It alleges violations of the Colorado Anti-Discrimination Act, the ADA and Title VII.

The request that came before the refusal is the part to read slowly. “I requested an accommodation: human-generated captioning for the interview. Unfortunately, Intuit did not provide me with this requested accommodation, instead saying that HireVue had built-in subtitles.” — D.K., writing for the ACLU, 7 April 2025.

That exchange is the whole problem in two sentences. A machine's handling of speech was offered as the equivalent of a person's. It was offered by a party with no way of knowing whether it was. A promotion decision was then taken downstream of it. Nothing in that sequence requires opening the model.

The vendor's account of what its video interviews measure had already been challenged six years earlier. A separate complaint quoted HireVue on candidates' “intonation,” “inflection,” and “emotions.” That is the same acoustic surface a Deaf applicant asking for human captioning had every reason to distrust.

  • The ground under the decision is wide: language and accent recognition, emotion and affect labels, age and state proxies, the ambiguity in the annotations themselves, fairness, and where appropriate use stops.
  • The failure that arrives first is performed feeling standing in for ordinary speech. IEMOCAP, the field's standard acted set, holds roughly 12 hours of audiovisual data from ten actors. A model fitted on that material is afterwards read as a description of how people sound at work.
  • The evidence anyone should have asked for is construct validity against an appropriate reference measure. Report it beside a breakdown by speaker group. That is the kind of breakdown that produced an average word error rate of 0.35 against 0.19 across five commercial ASR systems.
  • The practical response is to retire the broad labels these systems are sold under. The 2019 FTC complaint named four: “cognitive ability,” “psychological traits,” “emotional intelligence,” and “social aptitudes.” Replace them with something observable and bounded in time.

Case

A 68-page review of facial movements, and a prohibition in force since 2 February 2025

Facial movements are not reliable, specific or generalizable indicators of emotion. That is the conclusion of a 68-page review published in 2019 in Psychological Science in the Public Interest. The Association for Psychological Science commissioned it. Lisa Feldman Barrett and four co-authors wrote it. The sentence underneath the conclusion is the one worth carrying: “Yet how people communicate anger, disgust, fear, happiness, sadness, and surprise varies substantially across cultures, situations, and even across people within a single situation.” — Barrett and colleagues, 2019.

Be exact about the scope, because the temptation here is to smuggle. The review's evidence is facial movements, not vocal ones. It does not measure prosody, and it does not license a conclusion about prosody by itself. What it does establish is the failure of the assumption the whole enterprise runs on. That assumption: a category of feeling maps onto an observable behaviour in a stable way, across people, cultures and situations. It failed when it was tested. It failed in the modality with the largest literature and the cleanest measurement. Anyone extending that to voice has to argue for the extension out loud rather than inherit it.

The European Union then wrote a limit into law without waiting for the argument. Article 5 of the AI Act prohibits AI systems that infer emotions in the workplace and in education institutions, except where the purpose is medical or safety-related. The prohibition has applied since 2 February 2025. The two bind in different ways. The review says the reading was not supported by the evidence in the modality studied most. The regulation says that in the one setting D.K.'s interview took place in, it may not be attempted at all.

Analogy

Reading weather from one window

Set the review and the regulation aside for a moment and the problem is still visible from the inside. A view through one window at one moment is a poor basis for a claim about the climate of a region. It is roughly the basis a paralinguistic score rests on. The observation contains clues. It also contains local conditions and selection effects. The room. The handset. The reason this particular recording was the one that got made. And, in a recorded interview, whether the audio reaching the scorer had been transcribed by a person or by a second machine.

The comparison is generous to the score in one respect. Weather is not produced for an audience. Speech is, and it carries linguistic content on the same acoustics the model is reading.

A voice sample supports narrow observations, not sweeping claims about a person.

Voice is not a window into a mind

It is worth being exact, then, about what these models do estimate. Paralinguistic models estimate patterns beyond lexical content: language variety, speaking style, arousal, vocal effort, or affect as an annotator perceived it. Those labels are contextual. They are culturally dependent, variable over time, shaped by content and channel. What voice does not provide is a reliable general-purpose window into honesty, personality, intent, protected status, or private mental state.

A regulator went looking for an emotion AI that met data protection requirements and did not find one. On 26 October 2022 the UK Information Commissioner's Office warned organisations against deploying biometric emotion-analysis technology. Organisations failing to meet its expectations would be investigated. “Developments in the biometrics and emotion AI market are immature. They may not work yet, or indeed ever.” — Stephen Bonner, Deputy Commissioner at the ICO, 26 October 2022. The claim there is not that the products are imperfect. It is that nothing on the market cleared the bar.

Two separate questions follow from all this. Whether the score measures what it claims to measure is one. Whether it holds across languages, dialects, cultures and channels is the other. An aggregate validity result on the first says nothing about how a pipeline behaves for a Deaf applicant who asked for human-generated captioning and was told the product had subtitles. And the Barrett review is an attack on the first that no per-dialect breakdown would ever have surfaced. Neither answer on its own is enough to release such a system. High-stakes inferences require strong construct validity, consent, and a considered alternative. Often they should not be deployed at all.

Deploying a perception score as a fact about a person shifts the cost of being misread onto whoever happens to speak differently.

Visual

Whose judgment is the label recording, and which limit binds it

Both questions disappear the moment the system reports a single number. One score folds four things into each other: measurement, modeling, decision, and verification. Who decided what would count as “emotional intelligence”. Which model was fitted to that decision. What a manager is permitted to do with the output. How anyone would check it afterwards. All four arrive compressed into one figure on a dashboard, where none of them can be disagreed with separately.

The path below keeps the four choices where they can be argued with. The order matters. The earliest of them is the one nobody goes back to. The last of them is not always a limit you set for yourself. Two legal limits already reach this exact setting, and they are shaped differently. The AI Act prohibits workplace emotion inference outright, outside medical or safety purposes. The Illinois Artificial Intelligence Video Interview Act, in force since 1 January 2020, compels disclosure instead. It covers an employer that asks applicants to record video interviews, uses AI analysis of those videos, and is considering applicants for positions based in Illinois. That employer must “Notify each applicant before the interview that artificial intelligence may be used to analyze the applicant's video interview and consider the applicant's fitness for the position.” — Section 5 of the Artificial Intelligence Video Interview Act.

The Act asks for three further things. An explanation of how the AI works and what general types of characteristics it uses to evaluate applicants. The applicant's consent, before the interview. Deletion of the applicant's interviews within 30 days of a request. Read the first of those against the first layer of the path. A team that cannot state the construct cannot explain what general types of characteristics the system uses. So the disclosure obligation lands on the definition step, not on the compliance step. And a vendor that answers it with “psychological traits” has disclosed a label, not a characteristic.

FigureLayers · 4 layers
  1. 01

    Define the construct

    State whether the target is acoustic state, self-report, observer judgment, diagnosis, or operational outcome.

  2. 02

    Design annotation context

    Specify language, culture, task, time window, and rater information.

  3. 03

    Separate content and channel

    Control lexical leakage, device, room, and script correlations.

  4. 04

    Set governance limits

    Restrict use, expose uncertainty, audit disparities, and provide recourse.

Governance limits are drawn around the construct you defined, so a construct that was wrong from the start gets fenced in rather than questioned.

Key idea

Twelve acted hours, four hundred natural ones

Claims about what a voice reveals come apart in specific ways rather than all at once. A broad claim survives only until one of them applies. These four are where the damage usually starts in an employment setting. 1) Using acted emotions as evidence for natural workplace behavior. 2) Accent classification enabling discrimination or surveillance. 3) Lexical content leaking the label into an “acoustic” model. 4) Managers treating probabilistic perceptions as objective employee traits.

The first one is measurable in the corpora themselves, which is why it should never have to be argued abstractly. IEMOCAP, released in 2008 and still the field's standard acted set, holds roughly 12 hours of audiovisual data from ten actors performing scripts and improvisations. The MSP-Podcast corpus is assembled the other way round: “The corpus consists of over 400 hours of diverse audio samples from various audio-sharing websites, all of which have Common Licenses that permit the distribution of the corpus.” — from the MSP-Podcast Corpus abstract, 2025. Those samples are labelled by at least five raters each. Twelve hours of ten people performing on cue, against more than four hundred hours of people talking, is not a difference of degree in a training set. It is a different object. A number learned from the first gets quoted about the second every time the substitution goes unremarked.

The fourth failure mode has a documented instance with a regulator attached to it. On 6 November 2019 the Electronic Privacy Information Center asked the Federal Trade Commission to investigate HireVue. The company's video-interview assessments were unproven, EPIC argued, and constituted unfair and deceptive trade practices under Section 5 of the FTC Act. HireVue had said those assessments drew on candidates' “intonation,” “inflection,” and “emotions.” EPIC's own words: “HireVue has failed to demonstrate any legitimate purpose for the collection of job candidates’ biometric data or for the use of secret, unproven algorithms to assess the “cognitive ability,” “psychological traits,” “emotional intelligence,” and “social aptitudes” of job candidates.” — EPIC, complaint filed with the Federal Trade Commission, 6 November 2019. Note what those four names are. They are not descriptions of an acoustic measurement. They are properties of a person. A manager who reads a score sold under one of them has already been told what it means before seeing a single figure. The fourth failure mode is complete at that moment rather than later. Before such a score reaches anyone it has to be shown to measure the construct it names, gathered with consent, and weighed against a non-vocal alternative that does the same job.

Ten performers, about twelve hours, an emotion on cue. An employment decision turns on how someone speaks when nobody asked them to perform, and accuracy on the first does not carry over to the second.

Comparison

Even the checkable level ships with measurable label error

Which of the four you are exposed to depends on how far the system reaches. Language identification estimates the spoken language or variety under a defined inventory — a closed question with a checkable answer. Acoustic state estimation and trait inference reach a great deal further, and the evidence each one owes grows with the reach.

It is worth pricing the closed question before treating it as free. VoxLingua107 is a spoken-language-identification training set, scraped automatically from YouTube across 107 languages. Valk and Alumäe presented it in 2021. “The size of the resulting training set (VoxLingua107) is 6628 hours (62 hours per language on the average) and it is accompanied by an evaluation set of 1609 verified utterances.” — Valk and Alumäe, from the VoxLingua107 abstract. Post-filtering raised the proportion of correctly labelled segments to 98%. Read that from the other side: after the cleaning step, roughly two segments in every hundred still carry the wrong language label.

That is the level with the defined inventory, the verified evaluation set and an answer you can check by asking a speaker. Assessments built on intonation and inflection and then reported as “psychological traits” are engineered with the confidence appropriate to the first level and read with the ambition of the third. A measurement taken at one level settles nothing at another.

FigureComparison · 3 columns

Language identification

Estimates the spoken language or variety under a defined inventory.

  • Decision focus: Define the construct
  • Useful evidence: Construct validity against appropriate reference measures
  • Watch for: Using acted emotions as evidence for natural workplace behavior
  • Best used when its assumptions are documented for spoken language, accent, and paralinguistic analysis

Acoustic state estimation

Measures observable properties such as vocal effort or arousal proxies.

  • Decision focus: Design annotation context
  • Useful evidence: Cross-language, dialect, culture, and channel performance
  • Watch for: Accent classification enabling discrimination or surveillance
  • Best used when its assumptions are documented for spoken language, accent, and paralinguistic analysis

Trait inference

Claims a persistent personal property and usually requires far stronger evidence.

  • Decision focus: Separate content and channel
  • Useful evidence: Calibration and abstention on ambiguous speech
  • Watch for: Lexical content leaking the label into an “acoustic” model
  • Best used when its assumptions are documented for spoken language, accent, and paralinguistic analysis

Steps

Run a construct-validity review

Closing that distance is what a construct-validity review is for. It exists so that another team can challenge the leap from acted recordings to how people actually speak at work. That is the leap that was made silently in the case above. The challenge has to come while there is still time to refuse it. The record holds three things: what the stated construct assumes, one counterexample that breaks it, and what the governance limits trigger once that counterexample appears.

For an engagement score all three were available cheaply. The construct assumed that enthusiasm sounds one way. The counterexample was a calm regional speaker, scored low, behaving perfectly normally. And in a workplace the governance limit is not a matter of internal taste, because Article 5 has applied since February 2025.

FigureProcess · 4 steps
  1. 1. Name the claim precisely

    Replace broad labels such as “engagement” with observable and time-bounded definitions.

  2. 2. List alternative causes

    Consider words, task, microphone, health, culture, and environment.

  3. 3. Design falsification tests

    Shuffle text, change channel, and compare self-report with observer ratings.

  4. 4. Decide whether to stop

    Reject the use if the construct is vague, intrusive, or not actionable fairly.

A review that cannot name a precise construct, a non-intrusive way to measure it, and a fair action to take on the result has already produced its answer: refuse the use.

Example

0.35 against 0.19, and the same phrases in both columns

The review asks whether the construct is defensible at all. A report has to say how the system actually performs, and one number will not do it. Construct validity against appropriate reference measures tells you whether the score tracks anything real. It is the right place to start. What it cannot tell you is that recordings of performed feeling were standing in for how people sound at work. That substitution was made before the reference measure was ever chosen.

What a slice breakdown does tell you is on the record. Five commercial ASR systems — Amazon, Apple, Google, IBM, Microsoft — were tested on 19.8 hours of sociolinguistic interviews with 42 white and 73 black speakers. “We found that all five ASR systems exhibited substantial racial disparities, with an average word error rate (WER) of 0.35 for black speakers compared with 0.19 for white speakers.” — Koenecke and colleagues, PNAS, 23 March 2020. The decisive detail is the control: the gap persisted on a subset of identical phrases. Same words, same task, roughly twice the error rate. That is what separates what was said from how it was said. It is also the shape of finding that an aggregate accuracy figure is built to hide — the difference between a pattern and a complaint about one applicant.

So say what unit the score is in and who was rated. Say how uncertain it is and in what setting it was collected. Report performance across languages, dialects, cultures and channels beside it, on the model of a table with 0.35 in one row and 0.19 in another.

  • For the core task, construct validity against appropriate reference measures — the question the acted-versus-natural substitution answers before you get to ask it.
  • For what the system does in service, performance broken out by language, dialect, culture and channel. Add a matched-content control, of the kind that held the 0.35 against 0.19 gap in place across identical phrases.
  • For the robustness slice, calibration and abstention on ambiguous speech. How does the model behave when the audio does not support a confident answer? And is a request for human-generated captioning treated as an accommodation, or answered with built-in subtitles?
  • Over the working life of the system, downstream disparity, appeal outcomes and misuse incidents — the record that a complaint to a civil rights division would otherwise be the first sight of.

Report construct validity against appropriate reference measures together with the slice table that turns 0.35 against 0.19 into a finding — before anyone has to file for it. Report downstream disparity, appeal outcomes and misuse incidents beside them.

Key takeaways