Research
Multi-dataset Joint Pre-training of Emotional EEG Enables Generalizable Affective Computing
Overview Research area: Machine learning for electroencephalography (EEG)-based affective computing, specifically cross-dataset emotion recognition and EEG foundation-model pre-training. Technical lev

- arXiv
- 2510.22197
- Published
- 2025-10-25
- Authors
- Qingzhu Zhang, Jiani Zhong, Zongsheng Li, Xinke Shen, Quanying Liu
AI summary
Overview
- Research area: Machine learning for electroencephalography (EEG)-based affective computing, specifically cross-dataset emotion recognition and EEG foundation-model pre-training.
- Technical level: Advanced. The paper assumes familiarity with self-supervised pre-training, contrastive learning, covariance/statistical alignment, attention architectures, and EEG signal preprocessing.
- Scope: The paper introduces mdJPT, a label-free multi-dataset joint pre-training framework with two alignment losses and a hybrid encoder, evaluated for few-shot cross-dataset and zero-shot emotion recognition across six public EEG datasets.
What This Paper Is About
Most large EEG "foundation models" are pre-trained generically across many different tasks, which the authors argue dilutes the signals relevant to any one complex task such as emotion recognition. This paper instead pre-trains on several datasets that all target the same cognitive function (emotion) and uses statistical alignment to bridge the differences between those datasets. The goal is an emotion-decoding model that transfers to new subjects, new datasets, and even new emotion category definitions without extensive labels or per-subject calibration.
Key Contributions
- mdJPT (multi-dataset joint pre-training): A scalable pre-training framework for EEG-based emotion recognition that trains jointly on multiple emotion datasets and is validated against generic EEG foundation models in cross-dataset generalization.
- Cross-dataset alignment (CDA) loss: A loss that aligns second-order (covariance) statistics of latent EEG features across datasets and subjects, intended to mitigate inter-dataset and inter-subject distribution shifts and enable zero-shot generalization to new categories and unseen datasets.
- Inter-subject alignment (ISA) loss: A contrastive loss that treats EEG segments from two subjects viewing the same emotional stimulus as positive pairs and mismatched stimuli as negative pairs, allowing representation learning without explicit emotion labels and therefore across datasets with inconsistent label categories.
- Hybrid spatiotemporal encoder: An Mamba-like linear attention (MLLA) channel encoder combined with a spatiotemporal dynamics model using spatial transition convolutions and local attention, designed to capture long-term temporal dependencies and inter-channel dependencies in EEG.
Main Findings
- Few-shot cross-dataset classification: Under leave-one-dataset-out evaluation with a 1:3 subject split for classifier training/testing (repeated over 6 random splits), mdJPT achieved the best average scores across all five metrics, with average accuracy 56.22, precision 56.26, recall 56.01, F1 55.46, and AUROC 79.96. The improvements over the state of the art averaged 1.68%, 4.97%, 2.81%, 3.43%, and 4.57% absolute (3.08%, 9.69%, 5.28%, 6.59%, and 6.06% relative) for accuracy, precision, recall, F1, and AUROC respectively.
- Best per-dataset few-shot results: mdJPT obtained the best results on all metrics for SEED, SEED-V, and SEED-VII, and ranked in the top two on nearly all metrics for SEED-IV, FACED, and DEAP. On SEED-VII, the seven-class dataset, mdJPT reached 43.93 accuracy versus 30.29 for MMM, the next best reported method.
- Zero-shot generalization: Pre-trained on five datasets and evaluated directly on a held-out dataset with no fine-tuning, mdJPT outperformed the comparison methods on all datasets, with an average improvement of 11.9% (40.0% relative). On SEED, mdJPT exceeded the second-best model, EEGPT, by 17.05%. On the nine-category FACED task, other models performed close to chance level (11%) while only mdJPT performed better than chance. On DEAP, mdJPT reached 73.34% accuracy, which the authors note exceeded the best fine-tuned model.
- Scaling with more pre-training datasets: Using SEED-V as the target, performance improved consistently as more datasets were added to joint pre-training, and adding a dataset always helped relative to the setting without it. The maximum number of training datasets gave an 8.55% improvement (15.14% relative) over the best single-dataset configuration.
- Effect of the CDA loss weight: Accuracy was 63.52% with no CDA loss; small positive weights generally improved accuracy, with the best result, 65.02%, at a CDA factor of 0.02. Larger factors slightly reduced performance, which the authors attribute to overly strong alignment hindering emotion-discriminative learning.
- Effect of removing the ISA loss: Removing the ISA loss dropped accuracy to 30.51%, below the DE baseline of 45.58% on SEED-V, indicating it is critical both for inter-subject alignment and for learning emotion-related representations.
- MLLA versus Transformer: Replacing the MLLA channel encoder with a vanilla multi-head transformer reduced accuracy from 62.35% to 59.90%, precision from 62.91% to 60.39%, recall from 62.58% to 60.37%, and F1 from 62.49% to 59.73%; the transformer scored higher on AUROC (84.53 versus 79.38).
- Model compactness: mdJPT has fewer trainable parameters than existing pre-trained models, reported as 1.0M, and performance was stable across different random seeds.
Methodology in Plain English
The framework has three stages: pre-train an EEG encoder on several emotion datasets, optionally fine-tune a lightweight classifier on a few labeled subjects from the target dataset, then test on held-out subjects. In the zero-shot setting, the fine-tuning stage is skipped entirely and the pre-trained model is applied directly to a new dataset.
All EEG recordings, despite coming from different devices with different electrode layouts, are interpolated to a standardized 60-channel configuration based on the 10-20 International System. Each channel is split into overlapping patches and processed independently by the MLLA channel encoder, which uses an input gate, a linear attention module, and a forget gate to model long-range temporal dependencies efficiently. The resulting multichannel embeddings are then passed through a trainable spatial projection, spatial transition convolutions at multiple time scales, and local attention that estimates time-varying importance weights, producing spatiotemporal EEG patterns.
Two losses drive pre-training without emotion labels. The CDA loss computes each subject's covariance structure in the latent space, averages it into a subject-level centroid, and minimizes the pairwise Euclidean (Frobenius) distances between the centroids of all subjects in a batch, where each batch contains subjects drawn from multiple datasets. The ISA loss is contrastive: it pulls together representations of EEG segments from two subjects who saw the same stimulus and pushes apart segments from different stimuli, using a normalized temperature-scaled cross-entropy objective. The total pre-training loss combines ISA with a weighted CDA term.
For classification, encoder outputs are averaged over a 5-second window, concatenated across consecutive samples within a trial, smoothed with a linear dynamical system (LDS) model, and fed to a two-layer MLP classifier with ReLU activation and batch normalization. Training used the Adam optimizer, Python 3.12.3, PyTorch 2.3.1, and an NVIDIA GeForce RTX 3090 GPU. Comparison baselines were MMM, LaBraM, EEGPT (using publicly available pre-trained parameters) and a DE-feature baseline without pre-training.
Why This Matters
The paper argues that generic large-scale pre-training over heterogeneous EEG tasks is not the right fit for complex, nuanced targets like emotion, because task-relevant signals get diluted. Its results support task-specific multi-dataset pre-training as an alternative paradigm, and the finding that performance scales with the number of joint pre-training datasets offers a concrete recipe for the EEG community. Without labels or per-subject calibration, the method also addresses the practical reality that emotion datasets use different category definitions and have large inter-subject variability.
Real-world applications that follow from cross-dataset, calibration-free EEG emotion decoding:
- Emotion-aware brain-computer interfaces that adapt to a new user or a new recording setup without lengthy per-person calibration.
- Mental health and wellbeing monitoring, where affect-related brain signals could be tracked across heterogeneous clinical and consumer devices.
- Adaptive human-computer interaction, such as content, learning, or interface systems that respond to a user's affective state.
- Cross-site research collaboration, where models trained on one cohort's data can be reused on another cohort's data and equipment.
Industry relevance: the framework's compactness (1.0M trainable parameters) and its ability to transfer without fine-tuning matter for deployment on resource-constrained hardware. The paper notes that practical deployment is currently limited by the cumbersome nature of EEG hardware and suggests wearable devices could help overcome that barrier. The authors also flag risks of misuse in emotional surveillance or profiling without consent, and call for regulatory guidelines and privacy-preserving frameworks.
Future Directions
- Improving fine-grained emotion classification: Despite its advantages, mdJPT's performance on the nine-category FACED task is described as still not satisfactory, indicating residual feature distribution shifts that need further work.
- Broader emotion elicitation paradigms: Evaluation focuses on video-induced emotions; an exploratory analysis on the EmoEEG-MC dataset covers imagery-induced contexts, and the authors state that generalization to diverse elicitation paradigms requires further validation.
- Modeling individual differences: The approach addresses signal-level heterogeneity but does not model individual differences in emotional experience; the authors suggest incorporating personalized emotion ratings through soft contrastive learning.
- Demographic diversity, hardware, and ethics: The generalizability of the findings is constrained by limited demographic diversity (age, culture, health status) in existing EEG emotion datasets, and responsible deployment requires regulatory guidelines and privacy-preserving frameworks.
Target Audience
Researchers and engineers working on EEG foundation models, affective computing, and brain-computer interfaces, particularly those interested in transfer learning and domain adaptation across datasets and subjects. It is also relevant to practitioners who need label-efficient emotion decoding with small models, and to reviewers or students studying how task-specific pre-training compares with generic large-scale pre-training in neurophysiological machine learning.
Authors’ abstract
Task-specific pre-training is essential when task representations diverge from generic pre-training features. Existing task-general pre-training EEG models struggle with complex tasks like emotion recognition due to mismatches between task-specific features and broad pre-training approaches. This work aims to develop a task-specific multi-dataset joint pre-training framework for cross-dataset emotion recognition, tackling problems of large inter-dataset distribution shifts, inconsistent emotion category definitions, and substantial inter-subject variability. We introduce a cross-dataset covariance alignment loss to align second-order statistical properties across datasets, enabling robust generalization without the need for extensive labels or per-subject calibration. To capture the long-term dependency and complex dynamics of EEG, we propose a hybrid encoder combining a Mamba-like linear attention channel encoder and a spatiotemporal dynamics model. Our method outperforms state-of-the-art large-scale EEG models by an average of 4.57% in AUROC for few-shot emotion recognition and 11.92% in accuracy for zero-shot generalization to a new dataset. Performance scales with the increase of datasets used in pre-training. Multi-dataset joint pre-training achieves a performance gain of 8.55% over single-dataset training. This work provides a scalable framework for task-specific pre-training and highlights its benefit in generalizable affective computing. Our code is available at https://github.com/ncclab-sustech/mdJPT_nips2025.