Research
A Baseline Multimodal Approach to Emotion Recognition in Conversations
Overview Research area: Natural Language Processing and affective computing, specifically multimodal emotion recognition in conversations. Technical level: Intermediate. Scope: This preprint by Víctor

- arXiv
- 2602.00914
- Published
- 2026-01-31
- Authors
- Víctor Yeste, Rodrigo Rivas-Arévalo
AI summary
Overview
Research area: Natural Language Processing and affective computing, specifically multimodal emotion recognition in conversations. Technical level: Intermediate. Scope: This preprint by Víctor Yeste and Rodrigo Rivas-Arévalo documents a lightweight, reproducible baseline that combines transformer-based text classification and self-supervised speech representation models with late fusion on the SemEval-2024 Task 3 dataset built from the sitcom Friends.
What This Paper Is About
Emotion recognition in conversations aims to identify emotional states from utterances, where meaning can depend on context, tone, and conversational flow. The paper's goal is not to propose a novel state-of-the-art method but to provide an accessible reference implementation that combines text and audio cues and reports empirical behavior under a limited training protocol. It focuses on when a simple late-fusion ensemble improves over unimodal text or audio models.
Key Contributions
- Provides a lightweight multimodal baseline for emotion recognition in conversations using the SemEval-2024 Task 3 dataset from Friends, combining a transformer-based text classifier and a self-supervised speech representation model with a simple late-fusion ensemble.
- Documents an empirical comparison of four text models (RoBERTa, DistilBERT, DeBERTa, DistilRoBERTa) and three audio models (HuBERT, Wav2Vec2, Wav2Vec2-large-robust) under limited hyperparameter tuning and limited evaluation.
- Reports a late-fusion ensemble of Wav2Vec2 and RoBERTa achieving 62.97% accuracy, compared with the best unimodal text model RoBERTa at 50.68% and the best audio model Wav2Vec2 at 35.43%.
- Explicitly states limitations and scope, including reliance on accuracy and wall-clock execution time rather than Macro-F1 or multiple random seeds, and provides code and materials for transparency and future comparisons.
Main Findings
- Text models outperform audio models on accuracy. RoBERTa reached 50.68%, DistilBERT 49.82%, DeBERTa 48.29%, and DistilRoBERTa 41.83%.
- RoBERTa was the strongest text model. It achieved the highest text accuracy at 50.68%, likely due to optimized training and effective capture of contextual dependencies.
- DistilRoBERTa was fastest but least accurate. Its execution time was 5.61 seconds, compared with Ro
Authors’ abstract
We present a lightweight multimodal baseline for emotion recognition in conversations using the SemEval-2024 Task 3 dataset built from the sitcom Friends. The goal of this report is not to propose a novel state-of-the-art method, but to document an accessible reference implementation that combines (i) a transformer-based text classifier and (ii) a self-supervised speech representation model, with a simple late-fusion ensemble. We report the baseline setup and empirical results obtained under a limited training protocol, highlighting when multimodal fusion improves over unimodal models. This preprint is provided for transparency and to support future, more rigorous comparisons.