Research
MEG-to-MEG Transfer Learning and Cross-Task Speech/Silence Detection with Limited Data
Overview Research area: Non-invasive neural speech decoding from magnetoencephalography (MEG), transfer learning, and brain-computer interfaces (BCIs), sitting at the intersection of machine learning
- arXiv
- 2602.18253
- Published
- 2026-02-20
- Authors
- Xabier de Zuazo, Vincenzo Verbeni, Eva Navas, Ibon Saratxaga, Mathieu Bourguignon, Nicola Molinaro
AI summary
Overview
- Research area: Non-invasive neural speech decoding from magnetoencephalography (MEG), transfer learning, and brain-computer interfaces (BCIs), sitting at the intersection of machine learning and cognitive neuroscience.
- Technical level: Intermediate. The paper assumes familiarity with MEG sensor data, deep sequence models (Conformer), and standard decoding metrics (F1, balanced accuracy, AUC), but its central idea — pre-train on lots of data from one person, fine-tune on a few minutes from many others — is conceptually simple.
- Scope in one sentence: The paper shows that a Conformer-based MEG model pre-trained on 50 hours of single-subject listening data, then fine-tuned on 5 minutes per subject across 18 participants, improves both in-task and cross-task speech/silence detection and enables decoding between speech perception and speech production.
What This Paper Is About
Practical speech BCIs are limited because each person provides only minutes of usable neural data, yet current MEG decoders are typically trained from scratch per subject and per task. This paper asks whether a model can instead be pre-trained on a large single-subject MEG corpus and then adapted to new subjects with very little data, and whether such a model can generalize across different speech tasks — listening to speech, listening to one's own playback, and speaking aloud. The authors also test the deeper scientific question of whether perception and production share neural representations that a decoder can exploit.
Key Contributions
- First MEG-to-MEG transfer learning demonstration for a speech decoding task. A MEGConformer model is pre-trained on the LibriBrain dataset (over 50 hours of within-subject MEG from a single participant listening to English audiobooks) and fine-tuned on just 5 minutes per subject across 18 participants in the Bourguignon et al. dataset. The authors state this is the first successful application of transfer learning to a MEG speech decoding task (speech detection), noting prior MEG transfer work only used ImageNet-pretrained vision models for imagined speech.
- First cross-task MEG speech detection results. All six train-test pairings among Listen, Playback, and Production are evaluated, extending prior cross-task decoding work that had been limited to non-speech EEG/MEG.
- Evidence that production-trained models decode passive listening above chance. Models trained only on speech production decode listening and playback above chance, which the authors interpret as evidence that decoding relies on shared neural speech representations rather than task-specific motor activity alone.
- Lightweight, reusable fine-tuning adaptations. Three task-specific modifications are introduced:
RollAugment(a fast roll-based temporal augmentation that circularly shifts each training frame by 25%, 50%, and 75% of the window and concatenates the shifted copies), soft targets based on the fraction of speech within each window instead of hard 0/1 labels, and checkpoint selection by validation loss rather than F1-macro.
Main Findings
- In-task transfer learning helps, most clearly for listening. For the listening task, accuracy rose from 76.2 ± 4.8% (scratch) to 79.0 ± 4.8% (transfer), F1 from 85.5 ± 3.2% to 87.7 ± 3.2%, and AUC from 64.0 ± 9.4% to 68.7 ± 6.2% (W = 17.0, p = 0.005; the paper reports these as gains of +3.7% accuracy, +2.6% F1, and +7.3% AUC). Playback improved modestly and not significantly (W = 45.0, p = 0.163) and production much less (W = 61.0, p = 0.304), though the overall in-task effect across tasks and metrics was significant under a sign-flip permutation test (p < 0.001).
- Cross-task decoding works even without pre-training. Using models trained from scratch, all six cross-task combinations decoded above chance, each individually significant (p < 0.05) and jointly significant (p < 0.001). Accuracy ranged from 65.0% to 73.4%; perception-to-perception transfer was strongest (listen-to-playback 72.5 ± 3.6%, playback-to-listen 73.4 ± 2.9%) and production-to-perception weakest (production-to-listen 66.1 ± 6.3%, production-to-playback 65.0 ± 5.9%).
- Transfer learning gives its largest benefit in cross-task settings. The abstract reports in-task accuracy gains of 1–4% versus larger cross-task gains of up to 5–6%. Listen-to-playback improved by +6.1% accuracy, +4.2% F1, and +2.9% AUC (W = 3.0, p < 0.001); playback-to-listen improved by +6.3% accuracy, +4.1% F1, and +3.7% AUC (W = 3.0, p < 0.001); listen-to-production improved by +5.3% accuracy and +3.6% F1 (W = 22.0, p = 0.016); production-to-listen gained +4.8% accuracy, +3.1% F1, and +3.7% AUC (W = 36.0, p = 0.048); production-to-playback improved by +5.1% accuracy, +3.3% F1, and +3.7% AUC (W = 33.0, p = 0.048). The overall cross-task benefit was significant (p < 0.001).
- Transfer effects are larger for cross-task than in-task decoding in effect-size terms. Figure 1 shows in-task F1 improvements of 0.5–2.2% versus cross-task improvements of 1.7–3.5%, with cross-task conditions involving production showing greater variability across subjects.
- Bidirectional transfer is asymmetric. Listening and playback transferred in both directions with similar performance (listen-to-playback 86.5% F1; playback-to-listen 87.1% F1), but perception-to-production transfer (listen-to-production 85.3% F1; playback-to-production 83.7% F1) substantially outperformed production-to-perception transfer (production-to-listen 80.1% F1; production-to-playback 79.0% F1). The authors attribute this to production engaging additional motor planning, efference copy, and somatosensory processes absent in passive perception.
- Production models decode passive listening above chance. Despite the asymmetry, production-trained models still decoded listening and playback reliably, which the authors treat as evidence for shared neural circuitry between perception and production, consistent with dual-stream models of speech processing.
- Subject-level responses are heterogeneous. In perception tasks, 15 of 18 subjects improved, while 2 subjects (identified as 1 and 2) showed clear negative effects. In production, 16 of 18 subjects benefited, but Subject 16 showed a marked negative effect of -13.3% F1. The authors interpret this variability as motivation for subject-adaptive approaches in practical BCI applications.
Methodology in Plain English
The authors combined two MEG datasets with deliberately contrasting properties. For pre-training they used LibriBrain, a public dataset with over 50 hours of MEG from one right-handed male native English speaker listening to Sherlock Holmes audiobooks, recorded on a 306-channel Elekta/MEGIN system (102 magnetometers, 204 planar gradiometers) and processed with the LibriBrain Competition pipeline, downsampled to 250 Hz. Roughly 76.7% of labeled time is speech. For fine-tuning and evaluation they used Bourguignon et al., a non-public dataset of 18 healthy adult native Spanish speakers (9 female, 8 male, 1 unreported; mean age 23.9 years) recorded on a 306-channel Elekta/MEGIN system in a different lab and also downsampled to
Authors’ abstract
Data-efficient neural decoding is a central challenge for speech brain-computer interfaces. We present the first demonstration of transfer learning and cross-task decoding for MEG-based speech models spanning perception and production. We pre-train a Conformer-based model on 50 hours of single-subject listening data and fine-tune on just 5 minutes per subject across 18 participants. Transfer learning yields consistent improvements, with in-task accuracy gains of 1-4% and larger cross-task gains of up to 5-6%. Not only does pre-training improve performance within each task, but it also enables reliable cross-task decoding between perception and production. Critically, models trained on speech production decode passive listening above chance, confirming that learned representations reflect shared neural processes rather than task-specific motor activity.