Research
DementiaBank-Emotion: A Multi-Rater Emotion Annotation Corpus for Alzheimer's Disease Speech (Version 1.0)
Overview Research area: Natural language processing / speech emotion recognition applied to clinical populations — specifically Alzheimer's disease (AD) speech drawn from the DementiaBank Pitt Corpus.

- arXiv
- 2602.04247
- Published
- 2026-02-04
- Authors
- Cheonkam Jeong, Jessica Liao, Audrey Lu, Yutong Song, Christopher Rashidian, Donna Krogh, Erik Krogh, Mahkameh Rasouli, Jung-Ah Lee, Nikil Dutt, Lisa M Gibbs, David Sultzer, Julie Rousseau, Jocelyn Ludlow, Margaret Galvez, Alexander Nuth, Chet Khay, Sabine Brunswicker, Adeline Nyamathi
AI summary
Overview
Research area: Natural language processing / speech emotion recognition applied to clinical populations — specifically Alzheimer's disease (AD) speech drawn from the DementiaBank Pitt Corpus.
Technical level: Intermediate. The paper is primarily a corpus release and annotation-methodology paper, with an exploratory acoustic analysis component using standard speech features (eGeMAPS) and linear mixed-effects models.
One-sentence scope: The paper introduces DementiaBank-Emotion v1.0, the first multi-rater emotion annotation corpus for Alzheimer's disease speech, covering 1,492 participant utterances from 108 speakers labeled for Ekman's six basic emotions plus neutral, and reports that AD patients express significantly more non-neutral emotions than controls while showing attenuated prosodic differentiation.
What This Paper Is About
Computational work on Alzheimer's disease speech has focused almost entirely on cognitive-linguistic markers (lexical diversity, syntax, pauses) and has largely ignored how emotion is expressed. Existing speech emotion datasets such as IEMOCAP, MSP-IMPROV, and RAVDESS are built on healthy speakers using acted or scripted speech, so they do not capture the atypical prosody, pragmatic functions, and clinical ambiguities of AD speech. This paper addresses that gap by building and releasing a multi-rater, expert-annotated emotion corpus for AD speech, documenting the annotation challenges, and running exploratory acoustic analyses on the resulting labels.
Key Contributions
- The release of the first emotion-annotated corpus for AD speech: 1,492 utterances from 108 speakers with multi-rater labels (Fleiss' κ = 0.23–0.31 post-calibration), drawn from the ADReSS 2020 Challenge training set (54 AD and 54 control speakers, matched for age and gender).
- Documentation of annotation challenges specific to clinical speech, together with calibration workshop materials addressing issues such as distinguishing face-saving laughter from genuine joy.
- The finding that AD patients express more non-neutral emotions than controls (16.9% vs. 5.7%), with exploratory evidence suggesting reduced prosodic differentiation (termed "acoustic flattening") that the authors state requires replication.
- Evidence that loudness differentiates emotion categories within AD speech, suggesting partially preserved emotion–prosody mappings.
Main Findings
- Higher non-neutral emotion rate in AD: AD patients produced non-neutral emotions at 16.9% of 615 valid labeled utterances versus 5.7% of 731 control utterances (χ²(1) = 38.45, p < .001). Neutral was 83.1% of AD labels and 94.3% of control labels.
- Joy and surprise dominate non-neutral labels: Joy was the most frequent non-neutral emotion in both groups (AD: 7.6%, 47 utterances; control: 3.3%, 24 utterances), followed by surprise (AD: 4.2%, 26 utterances; control: 0.8%, 6 utterances). Sadness was 2.4% (15) in AD versus 0.7% (5) in controls; anger 1.3% (8) versus 0.1% (1); disgust 1.0% (6) versus 0.8% (6); fear 0.3% (2) versus 0.0% (0).
- Acoustic flattening (exploratory, limited sample): A significant Group × Sadness interaction emerged for F0 (β = −5.20, SE = 2.28, p = .023). Control speakers showed a substantial F0 decrease relative to their neutral baseline when expressing sadness (Δ = −3.45 semitones), whereas AD speakers showed virtually no change (Δ = +0.11 semitones). Bootstrap resampling at the speaker level (1,000 iterations) gave a 95% confidence interval of [−5.70, −0.22]. The authors caution that this rests on sadness samples of n = 5 utterances from 3 control speakers and n = 15 utterances from 11 AD speakers, and requires replication.
- Loudness differentiates emotions within AD speech: One-way ANOVA showed a significant effect of emotion on loudness (F(4, 610) = 8.48, p < .001, η² = 0.053) but not on F0 (F(4, 610) = 2.11, p = .078). Post-hoc Tukey HSD showed joy and surprise had significantly higher loudness than neutral and sadness: joy vs. neutral (Δ = +0.59, p < .001), surprise vs. neutral (Δ = +0.64, p = .004), joy vs. sadness (Δ = +1.00, p = .002), and surprise vs. sadness (Δ = +1.04, p = .003).
- Linguistic group differences are marginal: AD patients produced marginally shorter utterances (MLU = 7.84 words) than controls (MLU = 8.70; t = −2.13, p = .035, d = −0.41). Total words per session and type-token ratio did not differ significantly.
- Voice quality largely null at the emotion level: AD patients showed significantly higher HNR at the speaker level (d = 0.42, p = .033), though the authors note this may reflect recording conditions rather than voice quality. Jitter, shimmer, HNR, and H1-H2 did not significantly differentiate emotions after speaker normalization. A similar group pattern to F0 was observed for shimmer (β = −0.53, p = .017).
- Ambiguity is concentrated in AD speech: Of 146 ambiguous utterances, 137 were AD and 9 control (18.2% of AD utterances versus 1.2% of control). Ambiguous AD utterances had significantly lower F0, F0 variance, loudness, and HNR, and fewer words.
- Reliability improved after calibration: For patient data, Fleiss' κ went from 0.094 (Batch 1, before the calibration workshop) to 0.313 (Batch 2), then 0.231 (Batch 3). Control data reached κ = 0.254 with 3 raters after excluding one rater (L3) due to divergent labeling.
Methodology in Plain English
The researchers took audio recordings and transcripts from the ADReSS 2020 Challenge training set — a cross-sectional subset of the longitudinal DementiaBank Pitt Corpus with matched AD and control groups of 54 speakers each. Speakers performed the Cookie Theft picture description task from the Boston Diagnostic Aphasia Examination. Only participant utterances were annotated for emotion, because investigator utterances are predominantly neutral prompts and backchannels.
Eleven raters with mixed expertise (nursing researchers, a nurse practitioner, a business professor, and computer science and polytechnic researchers) labeled utterances into seven categories: Ekman's six basic emotions plus neutral. Annotation was done in three rounds. Because the task is a picture description, annotators were instructed to prioritize prosodic delivery over lexical content — so a charged word like "robbing" with flat delivery would still be labeled neutral. Calibration workshops with a psychiatry professor and a memory care facility director resolved disagreements, distinguishing happy laughter from helpless laughter and treating dramatic pitch excursions or gasps as markers of surprise.
The final label for each utterance came from a hierarchical adjudication algorithm: a majority of at least half the raters decides the label; a tie between neutral and an emotion resolves to the emotion; remaining ties are broken by confidence-weighted scores; anything still unresolved is marked ambiguous (NaN). For control data, one rater's annotations were excluded and the majority threshold was 2 of 3.
Acoustic features were extracted with openSMILE using the eGeMAPS feature set, covering prosody (F0, loudness), voice quality (jitter, shimmer, HNR, H1-H2), and spectral parameters. F0 was converted to semitones relative to each speaker's mean to separate intonational dynamics from physiological baseline, and other features were z-scored within speaker. Statistical analysis used chi-square tests, one-way ANOVA with Tukey HSD post-hoc tests, and linear mixed-effects models with speaker as a random intercept.
Why This Matters
This is the first resource of its kind — a clinical-population emotion corpus with multi-rater labels — and it makes a case that emotion in AD speech is not simply a degraded version of healthy emotional expression but has its own pragmatic functions, such as laughter used to save face during word-finding difficulty. The paper argues this directly explains a discrepancy with prior automated work: Chou et al. (2025) found minimal emotion-related differences between AD and control groups using machine learning on acoustic features alone, whereas expert human annotators integrating prosodic, linguistic, and pragmatic cues found significantly more non-neutral emotion in AD speech.
Real-world applications:
- Training emotion-aware assistive technologies that reflect the actual emotional expressions of users with cognitive impairment, rather than patterns learned from acted healthy speech.
- Helping caregivers and clinicians interpret emotional cues from AD patients accurately — for example, not reading face-saving laughter as happiness.
- Informing the design of speech emotion recognition systems for clinical use, which may need multimodal cues rather than prosody alone.
- Supporting research into emotional expression as a possible complementary marker alongside existing linguistic and acoustic features for tracking disease progression.
Industry relevance: Speech emotion recognition is deployed in call centers, virtual assistants, and health-monitoring products. This paper shows that systems trained on healthy acted speech can systematically misclassify laughter and emotional cues in clinical populations, which matters for any product intended to serve older adults or people with cognitive impairment. The corpus also provides a realistic, difficult benchmark for testing whether SER models generalize beyond prototypical emotional expression.
Future Directions
- Replicate the acoustic flattening pattern with larger samples of non-neutral emotions — the sadness finding rests on only 5 control and 15 AD utterances — and investigate whether prosodic realization of emotion is genuinely attenuated in AD.
- Version 2.0 is planned to include the full DementiaBank longitudinal data, allowing examination of relationships between emotional expression and cognitive severity over time.
- Add phone-level segmentation through manually corrected forced alignment (the automatic attempt failed at roughly 20% due to AD disfluencies), enabling fine-grained analysis of phonation type, vowel space dynamics, formant trajectories, and discourse markers such as "uh" and "um."
- Run calibration workshops for control data as well, since the absence of a control-data workshop contributed to the much higher AD ambiguity rate (18.2% versus 1.2%) and limits comparability between groups.
Target Audience
Researchers building speech emotion recognition systems, clinical NLP practitioners working on dementia and cognitive decline, and speech-language researchers interested in the affective dimension of AD speech. The paper is also relevant to clinicians and care professionals who interpret emotional cues from AD patients, and to corpus linguists and annotation-methodology researchers interested in inter-rater reliability, calibration procedures, and the difficulty of labeling emotion in spontaneous clinical speech. The released annotation guidelines and calibration workshop materials make it particularly useful for teams planning their own clinical annotation efforts.
Authors’ abstract
We present DementiaBank-Emotion, the first multi-rater emotion annotation corpus for Alzheimer's disease (AD) speech. Annotating 1,492 utterances from 108 speakers for Ekman's six basic emotions and neutral, we find that AD patients express significantly more non-neutral emotions (16.9%) than healthy controls (5.7%; p < .001). Exploratory acoustic analysis suggests a possible dissociation: control speakers showed substantial F0 modulation for sadness (Delta = -3.45 semitones from baseline), whereas AD speakers showed minimal change (Delta = +0.11 semitones; interaction p = .023), though this finding is based on limited samples (sadness: n=5 control, n=15 AD) and requires replication. Within AD speech, loudness differentiates emotion categories, indicating partially preserved emotion-prosody mappings. We release the corpus, annotation guidelines, and calibration workshop materials to support research on emotion recognition in clinical populations.