Skip to content
AI.info

Research

DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations

Overview Research area: Self-supervised representation learning ("foundation models") for intracranial electroencephalography (iEEG), spanning machine learning, clinical neurophysiology, and cognitive

DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations
arXiv
2512.19097
Published
2025-12-22
Authors
Danny Dongyeop Han, Yonghyeon Gwon, Ahhyun Lucy Lee, Taeyang Lee, Seong Jin Lee, Jubin Choi, Sebin Lee, Jihyun Bang, Seungju Lee, David Keetae Park, Shinjae Yoo, Chun Kee Chung, Jiook Cha

AI summary

Overview

  • Research area: Self-supervised representation learning ("foundation models") for intracranial electroencephalography (iEEG), spanning machine learning, clinical neurophysiology, and cognitive neuroscience decoding.
  • Technical level: Advanced. The paper assumes familiarity with Transformer attention, masked reconstruction pretraining, linear probing vs. finetuning, and scaling laws.
  • Scope: The paper introduces DIVER-1, a self-supervised iEEG foundation model for variable electrode layouts and timescales, pretrained on 5,310 hours of ECoG and SEEG, evaluated on two held-out benchmarks, and accompanied by a controlled compute-aware scaling study.

What This Paper Is About

Intracranial EEG offers direct, millisecond-scale recordings of human neural activity, but recordings differ across patients and centers in channel count, anatomical coverage, electrode type, referencing scheme, and recording protocol, so reusable representations are hard to learn. Existing iEEG foundation models are limited by restricted attention designs, fixed pretraining input interfaces, and small pretraining corpora, and prior work found that simple spectrogram-based linear decoders often beat pretrained iEEG models under leakage-aware evaluation. This paper's goal is an iEEG-specific architecture and input interface that transfers across subjects, tasks, and centers, plus the first controlled scaling study for self-supervised iEEG pretraining to determine which axes — data, parameters, training duration, or subject diversity — actually drive downstream performance.

Key Contributions

  1. A new architecture for variable electrode–time inputs. DIVER-1 combines any-variate electrode–time attention (full self-attention over all C·N electrode–time tokens, with RoPE for temporal offsets and a binary same-channel versus cross-channel attention bias), Spatio-Temporal Resampling (STR), Spatio-Temporal Conditional Positional Embedding (STCPE), and a Multi-Domain Reconstruction Objective (MDRO). Two patch-scale variants are instantiated: DIVER-1-0.1s and DIVER-1-1s.

  2. Large-scale cross-dataset pretraining and transfer. Pretraining on 5,310 hours of iEEG (352k channel-hours, 37 subjects) — roughly 54 times the BrainTreeBank-based pretraining volume — with transfer evaluated across 16 downstream evaluations on the naturalistic Neuroprobe benchmark and the clinical MAYO seizure-detection benchmark.

  3. State-of-the-art results under cross-center shift. DIVER-1-0.1s achieves the best mean AUROC on Neuroprobe despite no BrainTreeBank pretraining, and is reported as the first iEEG foundation model to surpass the linear spectrogram baseline in mean AUROC. DIVER-1-1s achieves the top AUROC on MAYO seizure detection.

  4. The first controlled compute-aware scaling-law study for self-supervised iEEG pretraining (to the authors' knowledge), sweeping data scale, subject count, training duration, model size, and compute up to 1.83B parameters. Code is available (link given in the paper).

Main Findings

  • Neuroprobe (15 tasks, 1-s windows): DIVER-1-0.1s achieves the best mean AUROC among all evaluated pretrained iEEG models and supervised model baselines. On individual tasks it is state-of-the-art or competitive on most tasks, except Frame Brightness, where DIVER-1 and most baselines remain close to chance.
  • Cross-dataset shift is stricter for DIVER-1: BrainBERT, PopT, and BaRISTA are pretrained on BrainTreeBank, the same corpus used to build Neuroprobe, whereas DIVER-1 is pretrained on external adult iEEG from different recording centers spanning SEEG and ECoG, while Neuroprobe evaluates pediatric SEEG.
  • Linear probing is highly competitive: the frozen DIVER-1-0.1s outperforms its full-finetuning counterpart on Neuroprobe; on MAYO, linear probing remains nearly on par with full finetuning.
  • MAYO seizure detection (6-s windows): DIVER-1-1s achieves the strongest performance, substantially outperforming BrainBERT and Brant.
  • Variant–timescale matching: DIVER-1-0.1s is better matched to fine-grained 1-s naturalistic decoding, while DIVER-1-1s is better matched to the longer 6-s clinical seizure-detection context.
  • Architecture vs. data scale disentangled (Table 3): Under matched BrainTreeBank pretraining, DIVER-1-1s outperforms prior architectures on speech (0.770 ± 0.028 vs. BrainBERT 0.606 ± 0.021, PopT 0.657 ± 0.024, BaRISTA 0.678 ± 0.041) and onset (0.859 ± 0.018 vs. 0.754 ± 0.027, 0.648 ± 0.029, 0.745 ± 0.038). With the architecture fixed and external corpora scaled, a BTB-sized external corpus (Private 1.1%, x1.0 BTB) underperforms BTB pretraining (0.710 ± 0.015 speech, 0.758 ± 0.016 onset), but scaling to 17.7% (x15.8 BTB) and 70.6% (x63.3 BTB) closes and reverses the gap (0.840 ± 0.022 / 0.898 ± 0.014 and 0.847 ± 0.020 / 0.905 ± 0.013).
  • Ablations (Table 1, DIVER-1-0.1s, d_model = 256, 8 epochs): Removing resampling causes the largest drops — without channel and time resampling, speech falls to 0.807 ± 0.017, onset to 0.884 ± 0.011, volume to 0.628 ± 0.015, pitch to 0.552 ± 0.006, versus 0.890 ± 0.013, 0.922 ± 0.009, 0.698 ± 0.018, 0.572 ± 0.008 for DIVER-1-0.1s with 3-s window resampling. Vanilla attention (without RoPE and any-variate attention) reduces all four tasks (0.870, 0.898, 0.669, 0.560), while removing RoPE or the same/cross-channel bias alone has smaller, task-dependent effects.
  • Longer-context pretraining transfers to short windows: 3-s STR pretraining outperforms both 1-s pretraining without resampling and 1-s pretraining with STR.
  • Positional embedding flips under distribution shift: Removing the absolute 3D channel-position embedding improves all four ablated tasks under adult-to-pediatric transfer, but when pretraining and downstream evaluation are distribution-matched on BrainTreeBank, the 3D position embedding improves every task (Table 2: 0.770 ± 0.028, 0.859 ± 0.018, 0.639 ± 0.023, 0.558 ± 0.010 with it, versus 0.756 ± 0.023, 0.827 ± 0.018, 0.600 ± 0.024, 0.539 ± 0.011 without).
  • Other components: STCPE and spectral features provide modest gains, channel sub-modality has limited impact, and MDRO improves over raw-only reconstruction (0.875, 0.916, 0.680, 0.569).
  • Pretraining scaling: at fixed training duration, held-out loss decreases approximately linearly on log–log axes as compute, data size, and model size increase, consistent with Kaplan-style scaling. Larger models only outperform smaller ones after sufficient iterations; at short durations they can remain undertrained.
  • Downstream scaling favors data and duration over parameter count: dataset size gives the most reliable gains; model size shows weaker, more mixed effects. AUROC generally improves up to 32 training passes, with mild decline at 64. Subject diversity is non-monotonic under fixed total pretraining volume.
  • Compute-optimal allocation: with training fixed at two passes, DIVER-1-1s allocates much more of a compute increase to data growth and much less to parameter growth than LLM scaling in Kaplan et al.; with a fixed corpus, the efficient frontier favors small-to-mid-sized models trained longer over the largest models trained briefly. The recommended order is: increase unique data first, train sufficiently long second, scale parameters last.
  • Effective token volume: Brant's 281k channel-hour private SEEG corpus with 6 s patches yields 169M channel-patch tokens, whereas DIVER-1 yields 1.27B tokens for DIVER-1-1s and 12.7B tokens for DIVER-1-0.1s — 7.5× and 75× more training tokens at 1 s and 0.1 s resolution, respectively.

Methodology in Plain English

The researchers built an encoder that does not assume a fixed electrode montage. Each iEEG channel is split into non-overlapping temporal patches (P = 500 for DIVER-1-1s and P = 50 for DIVER-1-0.1s at 500 Hz), and a three-layer CNN turns each patch into a token. Three kinds of information are added to every token: a spatio-temporal conditional positional embedding (STCPE), which uses a sliding-window any-variate Transformer block (MOIRAI) so the bias is translation-equivariant in time and permutation-equivariant across channels, rather than order-sensitive as in channel-axis convolutions; channel metadata (sinusoidal MNI coordinate encodings following PopT, plus an electrode-type embedding for depth/grid/strip); and a spectral embedding from an FFT of each patch, following CBraMod. Tokens are then processed by any-variate Transformer blocks adapted from MOIRAI, where attention runs over all electrode–time tokens at once so that activity at one electrode and time can directly influence another electrode at a delayed time.

Pretraining is masked multi-domain reconstruction: random patches are masked and the model reconstructs them across complementary signal domains rather than only the raw waveform. To handle heterogeneous implants, STR samples random spatio-temporal subviews of each 30 s segment — C' ≤ min(C, 32) contacts and N' ≤ 30 temporal patches drawn from a scaled Beta(3, 1) distribution that favors larger subviews, with the temporal cap keeping the effective context budget comparable between the two variants. Because training the largest models required up to 128 A100 GPUs, the authors use maximal update parameterization (μP) so that hyperparameters tuned on small proxy models transfer to large ones via μTransfer, letting them scan 13M to 1.83B parameters under a single configuration. All recordings were high-pass filtered at 0.5 Hz, notch filtered at 60 Hz, resampled to 500 Hz, and cut into 30-second windows; clipping-based QA/QC removed electrodes when more than 3.33% of samples exceeded the clipping threshold and discarded segments when more than 50% of channels were affected; amplitudes were scaled from [-200, 200] µV to [-1, 1]. Evaluation used linear probing or full finetuning with a linear classifier on flattened tokens, against BrainBERT, PopT, Brant, and BaRISTA and benchmark-reported supervised baselines where available. Brant could not be evaluated on Neuroprobe because of its 6 s patch size, and PopT and BaRISTA could not be evaluated on MAYO because of missing channel coordinates.

Why This Matters

Impact on research. The paper reports that a pretrained iEEG model can beat the linear spectrogram decoder on Neuroprobe while training on no BrainTreeBank data, which reverses a prior result that had cast doubt on the value of iEEG foundation models. It also supplies the field's first controlled compute-aware scaling sweep for self-supervised iEEG pretraining, replacing heuristic claims that "bigger is better" with a data-constrained recipe.

Real-world applications:

  • Seizure detection from clinical iEEG monitoring, where DIVER-1-1s achieves the top AUROC on the MAYO benchmark using 6 s windows.
  • Naturalistic cognitive decoding from brief (1 s) windows of movie-watching iEEG, as measured by the 15 Neuroprobe tasks.
  • Brain–computer interfaces that must operate across patients with different electrode implants, since the encoder does not assume a fixed montage, channel count, or electrode sub-modality.
  • Precision neuromedicine and clinical monitoring more broadly, where electrodes remain implanted for days and produce long recordings during naturalistic hospital behavior.

Industry relevance. The results bear on how compute budgets should be spent for biomedical foundation models: the finding that data and training duration beat parameter count under fixed compute favors investing in data collection and longer training over the largest models, and the μP/μTransfer setup described (13M to 1.83B parameters, up to 128 A100 GPUs) offers a practical way to run hyperparameter scans without full-scale sweeps. The released code and the description of a multi-country, cross-center ECoG/SEEG corpus are relevant to groups building clinical neurotechnology products.

Future Directions

  • Extending to data scales beyond the 63.3× BTB external corpus to test whether the observed data-constrained regime persists and whether the crossover with distribution-matched pretraining continues to widen.
  • Reconciling the positional-embedding trade-off: absolute 3D channel coordinates help under matched distributions but hurt under adult-to-pediatric transfer, so the paper's own question of how to encode geometry across ages and populations remains open.
  • Resolving the subject-diversity result: subject count showed non-monotonic effects under a fixed pretraining volume, so the optimal balance between number of subjects and volume per subject is not settled.
  • Extending benchmark coverage: Brant could not be evaluated on Neuroprobe due to its 6 s patch size and PopT and BaRISTA could not be evaluated on MAYO due to missing channel coordinates, so comparable evaluation across all models remains incomplete.

Target Audience

Machine learning researchers working on foundation models and scaling laws for time series and neural data; computational and clinical neuroscientists analyzing iEEG, ECoG, or SEEG; and engineers building brain–computer interfaces or clinical seizure-monitoring and neuromedicine systems. Readers need grounding in Transformer architectures, self-supervised masked pretraining, and benchmarking practice to follow the methods and ablation sections in detail.

Authors’ abstract

Intracranial EEG (iEEG) provides direct, millisecond-scale recordings of human neural activity, but reusable representation learning is difficult because electrode layouts, anatomical coverage, referencing schemes, and recording conditions vary across patients and centers. We introduce DIVER-1, a self-supervised iEEG foundation model for variable-input recordings that combines any-variate electrode-time attention, spatio-temporal resampling, input-conditioned positional embeddings, and multi-domain masked reconstruction without assuming a fixed electrode montage. We pretrain two variants, DIVER-1-0.1s and DIVER-1-1s, on 5,310 hours of ECoG and SEEG spanning 352k channel-hours, roughly 54x the BrainTreeBank-based pretraining volume. We evaluate DIVER-1 on two held-out benchmarks: Neuroprobe for naturalistic cognitive decoding and MAYO for seizure detection. On leakage-aware Neuroprobe, DIVER-1-0.1s outperforms prior evaluated iEEG foundation models despite using no BrainTreeBank recordings, the corpus underlying Neuroprobe, during pretraining; it also exceeds the linear spectrogram decoder in mean AUROC and remains competitive with stronger nonlinear baselines, a level prior evaluated iEEG foundation models did not reach. DIVER-1-1s also achieves the top AUROC on MAYO seizure detection. Finally, we conduct, to our knowledge, the first controlled compute-aware scaling study for self-supervised iEEG pretraining, sweeping data scale, subject count, training duration, and model size up to 1.8B parameters. Our results indicate a data-constrained regime: expanding unique recordings and training sufficiently long are more reliable scaling axes than increasing parameter count alone. Code is available at link.

Read the original paper