Research
Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations
Overview Research area: Computer vision and multimodal affective computing, specifically automatic recognition of facial and vocal emotion expressions — here extended from single emotions to blended e

- arXiv
- 2601.13225
- Published
- 2026-01-19
- Authors
- Tim Lachmann, Alexandra Israelsson, Christina Tornberg, Teimuraz Saghinadze, Michal Balazia, Philipp Müller, Petri Laukka
AI summary
Overview
Research area: Computer vision and multimodal affective computing, specifically automatic recognition of facial and vocal emotion expressions — here extended from single emotions to blended emotions with relative salience.
Technical level: Intermediate. The paper is readable without deep math, but it assumes familiarity with pretrained video/audio encoders (CLIP, VideoMAE, HuBERT, WavLM), feature aggregation, and multi-label classification.
Scope: The paper introduces BlEmoRe, a 3,050-clip multimodal (video + audio) dataset of single and blended emotion portrayals with relative salience labels, and benchmarks existing state-of-the-art encoders on two tasks built from it — predicting which emotions are present, and predicting their relative prominence.
What This Paper Is About
People frequently feel more than one emotion at once — sadness mixed with anger, or happiness tinged with surprise — and the emotions in such blends are not necessarily equally strong. Most video-based emotion recognition systems, however, only recognize a single emotion per clip, and the few that attempt blends cannot say which emotion dominates. The authors argue this is largely a data problem, and they address it by releasing a new dataset with systematic salience annotations plus baseline benchmarks showing how far current models get.
Key Contributions
-
A new multimodal blended emotion dataset (BlEmoRe): 3,050 clips from 58 actors, covering six single emotions (anger, disgust, fear, happiness, sadness, neutral) and all 10 pairwise combinations of the five non-neutral emotions, each blend portrayed in three salience configurations (50/50, 70/30, 30/70). The dataset is split by actor into 43 training actors and 15 test actors, with five predefined cross-validation folds inside the training set.
-
The first publicly available blended emotion dataset with relative salience annotations: the authors state that none of the existing blended emotion datasets listed in their comparison table includes information on the relative salience of emotions within a blend. BlEmoRe contains 1,390 single-emotion recordings and 1,660 blended-emotion recordings.
-
Two reference evaluation metrics:
ACC_presence, which checks whether all present emotions are predicted with no false positives or false negatives, andACC_salience, which additionally requires the predicted proportions to reflect the correct ranking (equal salience, or one emotion dominant over the other). -
A benchmark of existing encoders: eight pretrained encoders (five video, two audio, one multimodal), early fusion combinations of top visual and audio models, three classifier configurations, and HiCMAE, evaluated with five-fold cross-validation and on a held-out test set, alongside trivial class-distribution baselines.
Main Findings
-
Blended emotion recognition is far from solved. On the held-out test set, the best presence accuracy was 0.332 (VideoMAEv2 + HuBERT) and the best salience accuracy was 0.180 (HiCMAE). Trivial baselines reach 0.074 presence and 0.033 salience on the test set.
-
Multimodal beats unimodal. On validation, the best single encoder was ImageBind (ACC_presence = 0.290, ACC_salience = 0.130), and the best audio encoder was WavLM (ACC_presence = 0.265, ACC_salience = 0.121). Early fusion combinations consistently outperformed unimodal variants, with ImageBind + WavLM achieving the best composite score (ACC_presence = 0.345, ACC_salience = 0.170).
-
Salience is hardest, and only one model handles it best. HiCMAE achieved the highest salience accuracy across all models, on both validation (ACC_salience = 0.180) and test (ACC_salience = 0.180), while reaching ACC_presence of 0.298 (validation) and 0.268 (test).
-
Feature aggregation outperformed subsampling. For spatiotemporal encoders, aggregating statistics across frames gave better results than processing fixed-length clips independently: VideoMAEv2 reached ACC_presence = 0.273 and ACC_salience = 0.106 with aggregation versus 0.260 and 0.124 with subsampling; VideoSwin reached 0.225 and 0.089 with aggregation versus 0.210 and 0.103 with subsampling.
-
Performance drops from validation to test, and rankings shift. The authors attribute this primarily to the post-processing thresholding step rather than training instability. Thresholds are tuned on validation folds, which introduces look-ahead bias there; on the test set they stay fixed, making results sensitive to small shifts in predicted distributions. The VideoMAEv2 aggregation model illustrates this — high test ACC_presence (0.293) but ACC_salience near trivial (0.054).
-
WavLM was the strongest unimodal model on the test set. It reached ACC_presence = 0.311 and ACC_salience = 0.084, ahead of the best visual model VideoMAEv2 (ACC_presence = 0.293, ACC_salience = 0.054).
-
Human judges also find blends difficult. On a subset of 18 actors, human judges achieved a mean accuracy of 0.43 for correctly identifying both emotions in a blend under multimodal (audio-visual) conditions, on a presence-based measure comparable to
ACC_presence. -
Task difficulty is roughly in line with MAFW. The authors note the best ACC_presence of 0.332 is roughly similar in scale to reported unweighted average recall values on MAFW, where AVF-MAE++ reported 17.25% UAR / 43.83% WAR and HiCMAE 13.29% UAR / 37.36% WAR on the 43-class compound-emotion task, though they caution the setups are not directly comparable.
-
PCA projections show visible structure. 2D projections of WavLM and VideoMAEv2 embeddings for happy, sad, and their blends show distinct clusters corresponding to single and blended emotion portrayals, after per-actor normalization applied for visualization only.
Methodology in Plain English
The stimuli were purpose-built rather than sorted into blend categories after the fact. Actors were instructed to portray single emotions and every pairwise blend of the five non-neutral emotions, in three salience conditions. The 70/30 and 30/70 labels do not mean actors produced exact proportional amounts — they indicate that one emotion was intended to be less prominent than the other. Actors expressed emotions through face and body and through non-linguistic vocalizations (cries, laughter, groans); no words were allowed. Recordings used studio lighting and dampened acoustics, a camera at approximately 1.2 m and a microphone 0.5 m above the actor directed at the chest.
For the machine learning experiments, the authors took pretrained encoders and used them as frozen feature extractors. Five video encoders were used (OpenFace 2.0, which yields 31 features per frame from 17 action unit intensities, 6 head pose parameters and 8 gaze features; CLIP; ImageBind; VideoMAEv2 ViT-B/16; and Video Swin Transformer), two audio encoders trained on raw 16 kHz waveforms (HuBERT LL-60k and WavLM Large, which output 1024-dimensional frame-level embeddings at 20 ms resolution), and the multimodal HiCMAE model, which used cropped and aligned faces, and whose fused embeddings were classified without architectural modification.
Frame-level and spatiotemporal embeddings were converted to fixed-size representations in two ways. The first computed seven statistics per feature dimension — mean, standard deviation, and the 10th, 25th, 50th, 75th and 90th percentiles — and concatenated them. The second, applied to spatiotemporal encoders, treated each 16-frame clip's embedding as a separate subsample and averaged logits at prediction time. All features were standardized using training-set statistics.
Labels were six-dimensional vectors, one dimension per emotion. Single emotions used one-hot vectors; blends used soft distributions reflecting salience (for example, 70% happiness and 30% sadness as [0, 0, 0, 0.7, 0.3, 0]).
Classification used a linear layer or a one-hidden-layer MLP with 256 or 512 ReLU units, with softmax over the six outputs. MLP-512 typically performed best, so results are reported for that configuration. Models minimized KL divergence between the predicted distribution and the target; HiCMAE used cross-entropy, which outperformed KL divergence in its validation experiments. Training used Adam with a batch size of 32 for aggregated features and 512 for subsampled features, a learning rate of 5 × 10⁻⁶, weight decay of 1 × 10⁻³, and up to 200 epochs (aggregation) or 300 epochs (subsampling); HiCMAE was fine-tuned for 50 and 100 epochs with a cosine schedule and linear warm-up, batch size 32. The number of epochs for the held-out test set was fixed at 100.
To turn soft outputs into discrete predictions, a grid search on the validation folds found two thresholds: α for which emotions count as present, and β for whether a blend counts as equal (50/50) or dominant/subdominant (70/30 or 30/70). For the test set, these values were carried over from the most successful validation fold and epoch. Model selection used a composite validation score, the average of presence and salience accuracy. Early fusion was implemented by concatenating aggregated features from paired visual and audio encoders — the top visual encoders (VideoMAEv2 and ImageBind) each paired with both audio encoders (HuBERT and WavLM).
Why This Matters
Impact on research. Blended emotion recognition has been held back by the absence of suitable data: prior datasets are either small (C-EXPR-DB with 400 samples, IMED with 285, CMED with 1,050), or heavily unbalanced across blend classes (MAFW, with 4,938 single and 4,058 blended clips across 32 blend classes). The authors describe BlEmoRe as the second largest dataset of blended emotion expressions and the most balanced one, and as the first with salience annotations. The paper also makes a methodological point: modeling blended emotion as multi-label soft classification with hard thresholds is fragile, and the drop from validation to test suggests alternative formulations deserve attention.
Real-world applications:
- Affect-aware human–computer interaction, where a system must respond appropriately to mixed feelings rather than a single detected emotion.
- Mental health and clinical monitoring, where ambivalent or mixed emotional states carry diagnostic meaning.
- Social robotics and virtual agents that need to gauge which of several emotions is dominant in an interaction.
- Content and media analysis, where audience reactions of mixed valence are poorly served by single-label classifiers.
Industry relevance. Video and audio emotion recognition is deployed in customer experience analytics, automotive driver monitoring, and interactive entertainment. Systems that can only output one emotion per clip misrepresent exactly the situations where accurate affect reading matters most. The paper's benchmark results — test accuracy in the 0.33 (presence) and 0.18 (salience) range — give a realistic picture of how much headroom remains before such capabilities are production-ready, and the released dataset gives teams a standard target to measure against.
Future Directions
- Move beyond thresholded soft classification. The authors suggest regression-based approaches that predict salience proportions directly, ranking-based formulations that capture relative prominence without hard thresholds, and multi-task setups that jointly optimize presence and salience. They also point to alternative training objectives such as bi-center loss, which anchors features to their base-emotion centers.
- Build models tailored to blends. The authors note that most state-of-the-art multimodal models (AVF-MAE++, T-MEP, HiCMAE, PTH-Net) were not specifically designed for blended emotion recognition, highlighting the need for further methodological advances.
- Improve transfer to unseen actors. Because the split is by actor, the reported results reflect generalization to new individuals; the validation-to-test drop and the shifting model rankings between the two sets remain open issues to resolve.
- Address the ecological and demographic limitations. The authors flag that expressions were acted under laboratory conditions, that the emotion set covers only five basic categories and their pairwise blends, that salience is reduced to three levels (50/50, 70/30, 30/70) from what is likely a continuous variable, and that the dataset has limited ethnic and cultural diversity as it was developed within a European context.
Target Audience
Researchers and practitioners in affective computing and multimodal machine learning who work on emotion recognition from face and voice, dataset builders interested in annotation schemes for compound or mixed affect, and applied engineers who need reference benchmarks for how well current encoders handle emotional blends. It is also relevant to psychologists and behavioral scientists studying blended emotion expression, given that the dataset design is grounded in emotion psychology and validated through human perception experiments.
Authors’ abstract
Humans often experience not just a single basic emotion at a time, but rather a blend of several emotions with varying salience. Despite the importance of such blended emotions, most video-based emotion recognition approaches are designed to recognize single emotions only. The few approaches that have attempted to recognize blended emotions typically cannot assess the relative salience of the emotions within a blend. This limitation largely stems from the lack of datasets containing a substantial number of blended emotion samples annotated with relative salience. To address this shortcoming, we introduce BLEMORE, a novel dataset for multimodal (video, audio) blended emotion recognition that includes information on the relative salience of each emotion within a blend. BLEMORE comprises over 3,000 clips from 58 actors, performing 6 basic emotions and 10 distinct blends, where each blend has 3 different salience configurations (50/50, 70/30, and 30/70). Using this dataset, we conduct extensive evaluations of state-of-the-art video classification approaches on two blended emotion prediction tasks: (1) predicting the presence of emotions in a given sample, and (2) predicting the relative salience of emotions in a blend. Our results show that unimodal classifiers achieve up to 29% presence accuracy and 13% salience accuracy on the validation set, while multimodal methods yield clear improvements, with ImageBind + WavLM reaching 35% presence accuracy and HiCMAE 18% salience accuracy. On the held-out test set, the best models achieve 33% presence accuracy (VideoMAEv2 + HuBERT) and 18% salience accuracy (HiCMAE). In sum, the BLEMORE dataset provides a valuable resource to advancing research on emotion recognition systems that account for the complexity and significance of blended emotion expressions.