Research
Functional embeddings enable Aggregation of multi-area SEEG recordings over subjects and sessions
Overview Research area: Machine learning for neural data — representation learning and transformer architectures applied to invasive human intracranial recordings (stereo-EEG / local field potentials)

- arXiv
- 2510.27090
- Published
- 2025-10-31
- Authors
- Sina Javadzadeh, Rahil Soroushmojdehi, S. Alireza Seyyed Mousavi, Mehrnaz Asadi, Sumiko Abe, Terence D. Sanger
AI summary
Overview
Research area: Machine learning for neural data — representation learning and transformer architectures applied to invasive human intracranial recordings (stereo-EEG / local field potentials).
Technical level: Advanced. The paper assumes familiarity with contrastive learning (Siamese networks, InfoNCE/SupCon), transformer encoder–decoder architectures, and intracranial electrophysiology.
Scope: The paper proposes a two-stage framework ("FunctionalMap") that learns subject-agnostic, region-structured embeddings of SEEG local field potentials and uses them as tokens for a transformer that aggregates recordings across subjects and sessions.
What This Paper Is About
Intracranial recordings are hard to combine across people because electrode count, placement, and covered brain regions differ widely, and standard anatomical alignment (such as MNI coordinates) does not guarantee functional similarity — two electrodes at matched coordinates can record different neural dynamics, or even different structures. The authors test whether neural signals can be aligned better by their functional characteristics than by anatomical coordinates, and whether such a representation can support a single shared transformer trained across many subjects without subject IDs or subject-specific heads.
Key Contributions
- Functional coordinates. A data-driven embedding that captures region-specific neural dynamics, giving each electrode a subject-agnostic "functional identity" intended to replace unreliable atlas-based alignment.
- Unified cross-subject modeling. Embedding-informed tokenization lets one transformer aggregate heterogeneous intracranial datasets — irregular electrode layouts, multiple regions — without per-subject models, subject-ID tokens, or fine-tuning.
- Empirical validation on a 20-subject multi-region LFP dataset. The functional embedding clustered regions and transferred zero-shot to held-out channels; in an 11-subject ablation of coordinate systems, functional embeddings significantly improved masked-region reconstruction compared with MNI-coordinate alignment.
- Open code release. All code is available at https://github.com/ICLR-Functional-Embedding/ICLR2026_Functional_Map.
The authors state that, to their knowledge, they are the first to learn a subject-invariant, region-structured functional coordinate system from invasive, multi-region human LFPs and use it for end-to-end training on a unified multi-session, multi-subject dataset.
Main Findings
- Simulation validates the embedding. On simulated data with parameterized region signatures across 10 subjects and 2 sessions (variability introduced by scaling, frequency shifts, and noise), the contrastively trained encoder (PSC) clustered recordings by region with high accuracy when trained on a subset of subjects and evaluated on held-out subjects.
- Embeddings are sensitive to physiologically relevant signal components. Perturbation-based attribution showed the smallest embedding shifts when freezing segments overlapping the bursts that defined a region's signature, and power spectral densities of signals nearest each region centroid matched the simulated signatures.
- The functional map is smooth. Pure sinusoidal inputs with gradually increasing frequency traced continuous trajectories through embedding space, passing near centroids of regions tuned to corresponding frequency bands.
- Within-subject encoding works. Per-subject encoders (PSC, 10 s segments) reached 75.78% ± 17.90% mean test accuracy (mean ± SD) on held-out time segments from electrodes seen during training across 20 subjects. Under stricter held-out-channel evaluation (regions with more than 3 channels), accuracy was 45.79% ± 18.44%, with some subjects reaching approximately 70%.
- Channel-level generalization is weaker than time-segment generalization. The authors attribute this to the limited number of electrodes per region within a subject (approximately 6–12), though average performance remained above chance.
- Joint training across subjects performs better. Trained jointly across the whole dataset, the encoder reached 80.71% ± 11.41% on unseen time segments and 49.18% ± 12.11% zero-shot on held-out channels — better than separate per-subject CNNs by roughly 5% on average for both evaluations, with reduced performance variability across subjects.
- The two contrastive objectives differ. PSC yielded slightly higher accuracy on held-out time segments, while MSC clearly outperformed on held-out channels, indicating stronger channel-robust generalization. MSC embeds on a hypersphere via cosine similarity; PSC operates in Euclidean feature space with more compact, centroid-like clusters.
- Functional tokens beat MNI tokens for reconstruction. In a masked-region reconstruction task where all VO channels were withheld and GPi/STN/VIM signals served as sources, three otherwise-identical transformers differed only in the channel coordinate system (MNI, Functional-1 = PSC, Functional-2 = MSC). Functional-2 significantly outperformed MNI (paired test on subject-level Pearson r, p ≈ 0.002); Functional-1 showed a positive trend that was smaller and not reliably significant.
- Functional coordinates yield more differentiated predictions. Because the four VO electrodes often shared nearly identical or uncertain MNI coordinates, the MNI model tended to generate similar waveforms for targets, whereas functional embeddings produced channel-specific predictions that better matched the heterogeneous ground-truth VO activity.
- The transformer outperformed per-subject baselines. Transformer + Functional Embedding-2 (FUNC2), trained on pooled multi-subject data, achieved the highest reconstruction accuracy in both left and right hemispheres, statistically significantly outperforming all per-subject baselines — a linear causal multi-output FIR, a Temporal Convolutional Network (TCN), a 2-layer GRU, a CopyBest correlation baseline, and a Zero predictor — under paired two-sided t-tests with Holm and BH/FDR correction (all corrected p < 0.001).
- No subject-specific machinery was used. All transformer results came from a single shared model with no subject-ID tokens, no subject-specific heads, and no fine-tuning.
Methodology in Plain English
Two-stage design. First, learn a shared "functional coordinate" for each electrode. Second, feed those coordinates into a transformer that learns relationships between brain regions.
Stage 1 — Functional embedding. The authors take short 10-second LFP segments from each recording channel and train a lightweight convolutional neural network (with batch normalization, GELU activations, and dropout) of the Siamese type, so that segments from the same brain region land close together in a 32-dimensional latent space, regardless of subject or session, while segments from different regions are pushed apart. Two contrastive objectives are compared: a pairwise Siamese contrastive loss with margin m = 0.5, and a modified supervised-contrastive (SupCon) approach using multi-positive InfoNCE with temperature τ = 0.2, plus a batchwise intra-class variance penalty weighted at λ_var = 0.05 and no projection head or view augmentations. Training pairs may come from one session or across subjects, sessions, and time points, so the encoder is forced to learn region signatures robust to subject and session variability.
Stage 2 — Functional transformer. For masked-region reconstruction, the model withholds every channel from one target region (VO) and must predict those waveforms from the remaining regions (GPi, STN, VIM). A 1D convolutional tokenizer turns each source channel into time-patch features; each channel's 32-D functional embedding is broadcast across its patches, fused with the convolutional features, and combined with a time-only positional code to form a source token. Masked targets have no signal, so query tokens are built from a learned per-patch query base fused with the target's functional embedding and the same positional code. A standard pre-LN encoder–decoder transformer performs self-attention over sources, then self-attention plus cross-attention over queries, and a linear head outputs waveform patches. Variable numbers of sources and targets are batched with padding masks. The training loss combines mean-squared error with a correlation term, MSE + λ(1 − ρ) with λ = 0.05, where ρ is the mean Pearson correlation across targets — added because pure MSE tends to favor amplitude shrinkage and flat predictions.
Data. Recordings come from 20 subjects with dystonia who underwent deep brain stimulation surgery, with SEEG electrodes implanted in basal ganglia and thalamic nuclei including GPi, STN, VO, VA, VIM, PPN, and SNr — over 442.86 electrode-hours of neural recordings, or 9.96 recording hours in total. Sessions consisted of self-paced finger-to-nose reaching or hand-squeeze movements, with blocks of roughly 60 s of active movement followed by roughly 30 s of rest, repeated at least four times per hand; analysis used the entire continuous recordings rather than trial- or state-segmented data. Only micro-contact LFPs were used (not macro-contacts). For transformer analyses, only 11 of the 20 subjects were used, based on MNI data availability; MNI coordinates were extracted for those 11 subjects, yielding approximately 600 macro-contacts.
Evaluation of the coordinate systems. The comparison held everything constant except the tokenizer's channel embedding (MNI coordinates, Functional-1 from PSC, or Functional-2 from MSC). Reconstruction quality was measured by Pearson correlation between predicted and true waveforms, computed per target channel and window and then aggregated per subject using Fisher-z averaging.
Why This Matters
Impact on research. The paper reframes cross-subject neural alignment as a representation-learning problem rather than an anatomical-registration problem. If validated more broadly, it provides a route to the large-scale, heterogeneous pretraining that has been difficult to achieve for invasive recordings, where electrode counts and coverage vary by clinical need rather than experimental design. It also speaks to a claimed gap between anatomical coordinates and functional identity — the authors note related work found no significant performance loss when ablating MNI positional encoding, consistent with their own result that MNI tokens are the weaker coordinate system.
Real-world applications (as suggested by the work's framing):
- Subject-agnostic pooling across irregular electrode montages, enabling multi-patient analysis without hand-built per-patient pipelines.
- Reusable pretraining across downstream tasks such as decoding, forecasting, and imputation of neural signals.
- Better-informed interpretation of deep-brain stimulation recordings in basal ganglia–thalamic circuits, where the target region and dynamics can differ between individuals.
- Foundations for cross-subject decoding systems in clinical electrophysiology where strict task structure and uniform sensor placement are unavailable.
Industry relevance. Neurotechnology and deep-brain stimulation device companies, brain–computer interface developers, and clinical neurophysiology groups all depend on models that transfer between patients. Lowering the requirement for uniform electrode grids and per-subject labels is directly relevant to scaling such systems, and the paper's approach deliberately avoids subject IDs and per-subject fine-tuning, which are costly in clinical deployment.
Future Directions
- Relax the reliance on region labels during contrastive training, exploring self-supervised and weakly supervised objectives that preserve functional locality at scale.
- Extend beyond basal ganglia–thalamic circuits to cortical ECoG and spiking activity to test whether the functional coordinate system generalizes.
- Run a fuller head-to-head comparison against population-level pretraining frameworks such as the Population Transformer (PopT) described in related work.
- Bridge the two directions — functional coordinates for channel identity and ensemble-level pretraining — which the authors suggest may yield the best of both worlds.
Target Audience
Researchers and engineers working on neural representation learning, intracranial electrophysiology (SEEG/LFP), and cross-subject or cross-session model transfer; machine learning practitioners interested in contrastive embedding methods and transformers for non-standard sensor sets; and clinical or translational groups in deep-brain stimulation and brain–computer interfaces who need models that generalize across patients. Readers without background in contrastive learning or invasive electrophysiology will find the method sections demanding.
Authors’ abstract
Aggregating intracranial recordings across subjects is challenging since electrode count, placement, and covered regions vary widely. Spatial normalization methods like MNI coordinates offer a shared anatomical reference, but often fail to capture true functional similarity, particularly when localization is imprecise; even at matched anatomical coordinates, the targeted brain region and underlying neural dynamics can differ substantially between individuals. We propose a scalable representation-learning framework that (i) learns a subject-agnostic functional identity for each electrode from multi-region local field potentials using a Siamese encoder with contrastive objectives, inducing an embedding geometry that is locality-sensitive to region-specific neural signatures, and (ii) tokenizes these embeddings for a transformer that models inter-regional relationships with a variable number of channels. We evaluate this framework on a 20-subject dataset spanning basal ganglia-thalamic regions collected during flexible rest/movement recording sessions with heterogeneous electrode layouts. The learned functional space supports accurate within-subject discrimination and forms clear, region-consistent clusters; it transfers zero-shot to unseen channels. The transformer, operating on functional tokens without subject-specific heads or supervision, captures cross-region dependencies and enables reconstruction of masked channels, providing a subject-agnostic backbone for downstream decoding. Together, these results indicate a path toward large-scale, cross-subject aggregation and pretraining for intracranial neural data where strict task structure and uniform sensor placement are unavailable.