Research
Scalable Diffusion Transformer for Conditional 4D fMRI Synthesis
Scalable Diffusion Transformer for Conditional 4D fMRI Synthesis Overview Research area: Generative deep learning applied to neuroimaging — specifically conditional synthesis of whole-brain, voxel-lev

- arXiv
- 2511.22870
- Published
- 2025-11-28
- Authors
- Jungwoo Seo, David Keetae Park, Shinjae Yoo, Jiook Cha
AI summary
Scalable Diffusion Transformer for Conditional 4D fMRI SynthesisOverview
Research area: Generative deep learning applied to neuroimaging — specifically conditional synthesis of whole-brain, voxel-level 4D task fMRI using latent diffusion transformers.
Technical level: Advanced. The paper assumes familiarity with denoising diffusion probabilistic models (DDPM), latent diffusion, VQ-GANs, transformers with adaptive normalization, classifier-free guidance, and standard fMRI analysis concepts such as GLM contrast maps, BOLD hemodynamics, and representational similarity analysis.
Scope: The paper presents the first conditional diffusion transformer for voxel-wise whole-brain 4D task-fMRI generation, combining 3D VQ-GAN latent compression with a CNN–Transformer hybrid backbone and joint task conditioning, evaluated on seven HCP task paradigms against a 3D U-Net diffusion baseline.
Venue context: arXiv:2511.22870v1 [cs.CV], 28 Nov 2025, submitted to the workshop "Foundation Models for the Brain and Body." License CC BY-NC-SA 4.0.
What This Paper Is About
Task fMRI measures how brain activity changes over time during cognitive tasks, but the data are extremely high-dimensional and vary enormously between people and scanning setups. The authors ask whether a modern generative model can synthesize realistic whole-brain 4D fMRI sequences that are conditioned on a specific cognitive task — producing not just plausible-looking images, but task-evoked activations and temporal dynamics that match real data. Previous work avoided this by generating only simplified outputs such as ROI time series, connectivity matrices, or static 3D activation maps; the paper states that no method has previously generated task-conditioned, whole-brain 4D fMRI with modern generative architectures.
Key Contributions
- First conditional 4D fMRI generative model: a voxel-level, whole-brain diffusion model conditioned on cognitive task labels, producing spatio-temporal brain dynamics rather than static maps or reduced representations.
- Scalable hybrid architecture: a latent diffusion transformer built from 3D VQ-GAN compression, a UNet-style CNN–Transformer hybrid backbone (convolutional residual blocks in early stages, transformer blocks with global attention in later stages), and joint conditioning through AdaLN-Zero, FiLM, and cross-attention. Ablation studies test each component.
- Neuroscience-aligned evaluation: instead of standard image metrics such as FID or Inception Score, the authors introduce three metrics — GLM activation map correlation, representational similarity analysis of inter-task structure, and condition specificity (Top-1 accuracy) — and report results on seven HCP task paradigms.
- Demonstration of scaling behavior: generative performance improves predictably as model capacity grows from tens to hundreds of millions of trainable parameters, and the model consistently outperforms a 3D U-Net (MONAI) diffusion baseline.
Main Findings
- Task-evoked maps are reproduced: On HCP task fMRI, the model reproduces task-evoked activation maps, with the abstract reporting a task-evoked map correlation of 0.83 and an RSA of 0.98, consistently surpassing the U-Net baseline on all metrics.
- Inter-task structure is preserved: RSA scores reach the theoretical maximum of 1.0, and condition specificity (Top-1 Accuracy) reaches perfect accuracy for 340 M parameter models and larger.
- Clear scaling trend: GLM activation map correlation, RSA, and condition specificity all improve as trainable parameters increase (38.1 M, 85.4 M, 151.5 M, 236.5 M, 340.3 M, and 462.9 M configurations). The authors note this mirrors scaling behavior seen in vision and language foundation models.
- The baseline lags despite comparable size: the MONAI 3D U-Net baseline, with comparable parameter count, is behind the proposed models on all three metrics.
- Hybrid backbone is best (Table 1 ablation): Hybrid (CNN early + Transformer mid/high), 236.5 M parameters, achieved Corr 0.7006, Top-1 0.8571, RSA 0.9526, versus All-CNN (235.0 M: 0.6289 / 0.7143 / 0.9195) and All-Transformer (238.0 M: 0.6734 / 0.7143 / 0.9448). The All-CNN model had the weakest task fidelity; the pure transformer was marginally better than All-CNN but still below the hybrid.
- Strong conditioning matters (Table 1 ablation): Full conditioning (AdaLN-Zero + Cross-Attn), 151.5 M parameters, reached Corr 0.6267, Top-1 0.5714, RSA 0.9207, while removing cross-attention (AdaLN-Zero only), 110.5 M parameters, dropped to Corr 0.5066, Top-1 0.7143, RSA 0.9001. Removing cross-attention reduced parameters but caused a notable performance drop.
- Task signals are weak relative to nuisance variance: A PCA analysis in Appendix A.1 shows that the largest variance component in HCP task fMRI is individual subject differences, followed by phase encoding direction (LR vs. RL), with task-evoked variability emerging only as a weaker source of variance.
- Temporal dynamics align with real BOLD responses: ROI time-series analysis using the Harvard–Oxford cortical and subcortical atlas resampled to 3 mm, with three representative ROIs per task, shows the model more faithfully reproduces temporal profiles than the MONAI baseline, which often underestimates or distorts condition-specific responses.
- Volume-wise compression suffices: Compressing volumes with a 3D VQ-GAN followed by latent diffusion supported effective 4D synthesis even without an explicit 4D compression network.
Methodology in Plain English
Diffusing noise directly over an entire 4D fMRI volume is computationally infeasible because the data are so large and training data are limited. The authors therefore work in a compressed latent space: a pretrained 3D VQ-GAN (fine-tuned on individual HCP task-fMRI volumes and then frozen) shrinks each volume into a compact representation with spatial dimensions divided by a factor of four while keeping the time axis.
Into that latent space they run a standard DDPM: noise is gradually added to the latent during training, and a neural network learns to predict the injected noise so that, at sampling time, it can start from pure noise and iteratively denoise down to a clean latent, which the VQ-GAN decoder turns back into a 4D fMRI volume.
The denoising network is the key design choice. It is arranged in a UNet-like hierarchy that mixes two kinds of blocks: convolutional residual blocks in the early stages, which carry a strong local, spatio-temporal inductive bias and train stably on limited data, and transformer blocks with global attention in later stages, which capture long-range dependencies across space and time and scale well. Features from different resolutions are fused by concatenation.
Task conditioning is injected in two complementary ways. Normalization-based conditioning modulates the network using the task label — AdaLN-Zero in the transformer blocks and FiLM in the convolutional residual blocks — while cross-attention lets the task embedding and latent tokens exchange information directly. Class dropout during training enables classifier-free guidance at sampling.
Training used AdamW with 400k steps, a linear diffusion noise schedule, batch size 16, EMA for sampling, T = 1000 diffusion steps (β_start = 0.0015, β_end = 0.0195), class dropout rate 0.05, weight decay 0.01, β1 = 0.9, β2 = 0.99, and EMA decay 0.9999. Everything ran on a single NVIDIA A100 40GB GPU with bfloat16 automatic mixed precision. Temporal frames were stacked along the channel dimension for joint spatio-temporal modeling in both the proposed models and the baseline.
Data came from the Human Connectome Project task-fMRI dataset (minimally preprocessed), with seven paradigms — working memory, emotion, language, motor, relational, social, and gambling — each reduced to one representative condition: 2-back places, fear, loss, story, relation, right hand, and mental, respectively. Preprocessing downsampled to 3 × 3 × 3 mm³, resampled to TR = 1.44 s, and cropped background voxels. Treating each condition block as an instance and extracting approximately 18 s (about 12 TRs) from onset yielded 34,632 instances from 1,083 participants, split by subject 90/5/5 into 31,168 training instances from 975 subjects.
Evaluation uses first-level GLMs per subject and second-level random-effects GLMs at the group level to produce z-maps. Real-data GLMs were fit on the held-out test set; for synthetic data, 100 samples per condition were generated and group-level GLM maps computed the same way. Fidelity is the voxelwise Pearson correlation between real and synthetic group maps. RSA compares the off-diagonal entries of K × K representational dissimilarity matrices (Pearson correlation between group-level activation maps of each pair of tasks) for real versus synthetic data. Condition specificity ranks each generated sample's voxelwise correlations against all real group-level maps and measures how often the correct task wins.
The baseline is a 3D U-Net conditional diffusion model from the MONAI generative package, adapted for 4D fMRI using the same channel-stacking strategy and the same VQ-GAN latent space, in configurations from 41.3 M to 475.8 M parameters. The authors note that a prior α-GAN study for fMRI generation could not be replicated because its implementation is not publicly available.
Why This Matters
Impact on research. The paper argues that generative neuroimaging may benefit from scaling in the way vision and language foundation models have, and that diffusion transformers exhibit clear, predictable scaling laws on this task. It also argues that the best design is not the purest one: a hybrid CNN–Transformer backbone beat both an all-convolutional and an all-transformer variant, and strong conditioning (both normalization-based and cross-attention-based) was necessary to capture subtle task-specific signals buried under subject and acquisition variability. The three proposed metrics offer a neuroscience-grounded alternative to FID-style image scores for judging whether synthetic brain data are actually task-valid.
Potential applications named in the paper:
- Virtual (in-silico) experiments that probe brain dynamics under controlled cognitive manipulations.
- Cross-site harmonization of fMRI data.
- Principled data augmentation for downstream neuroimaging models.
- Simulation-based inference and, in the longer term, precision psychiatry.
Industry relevance. Building foundation-model-style generative systems for brain data has implications for clinical trial simulation, neuroimaging software pipelines, and any setting where acquiring large volumes of task fMRI is expensive. The reliance on a single A100 40GB GPU for training keeps the approach within reach of well-resourced academic and industry labs.
Future Directions
- Training across larger datasets: the authors state that future work will explore training on larger data collections to test whether the observed scaling trend continues.
- Integration of multimodal signals: combining fMRI generation with other data modalities is listed as a next step.
- Downstream applications: virtual experiments, cross-site harmonization, simulation-based inference, and precision psychiatry are named as the intended destinations, none of which are demonstrated in this paper.
- Better 4D compression: the authors note that volume-wise compression with a 3D VQ-GAN served as a practical stand-in because a high-quality 4D compression network is not available; developing one is a natural open problem.
Target Audience
Researchers at the intersection of generative modeling and neuroimaging — particularly those building diffusion or transformer-based models for brain data, and cognitive neuroscientists interested in what synthetic task fMRI can and cannot reproduce. The paper is also relevant to practitioners who need augmentation or harmonization tools for fMRI datasets, and to readers tracking whether foundation-model scaling behavior transfers outside vision and language. It is less suited to readers without a background in diffusion models and fMRI analysis, since the key arguments rest on latent diffusion mechanics and GLM/RSA evaluation.
Authors’ abstract
Generating whole-brain 4D fMRI sequences conditioned on cognitive tasks remains challenging due to the high-dimensional, heterogeneous BOLD dynamics across subjects/acquisitions and the lack of neuroscience-grounded validation. We introduce the first diffusion transformer for voxelwise 4D fMRI conditional generation, combining 3D VQ-GAN latent compression with a CNN-Transformer backbone and strong task conditioning via AdaLN-Zero and cross-attention. On HCP task fMRI, our model reproduces task-evoked activation maps, preserves the inter-task representational structure observed in real data (RSA), achieves perfect condition specificity, and aligns ROI time-courses with canonical hemodynamic responses. Performance improves predictably with scale, reaching task-evoked map correlation of 0.83 and RSA of 0.98, consistently surpassing a U-Net baseline on all metrics. By coupling latent diffusion with a scalable backbone and strong conditioning, this work establishes a practical path to conditional 4D fMRI synthesis, paving the way for future applications such as virtual experiments, cross-site harmonization, and principled augmentation for downstream neuroimaging models.