Research
An explainable framework for the relationship between dementia and glucose metabolism patterns
Summary: An Explainable Framework for the Relationship between Dementia and Metabolism Patterns Overview Research area: Machine learning for neuroimaging — specifically, generative deep learning (vari

- arXiv
- 2601.20480
- Published
- 2026-01-28
- Authors
- C. Vázquez-García, F. J. Martínez-Murcia, F. Segovia Román, A. Forte, J. Ramírez, I. Illán, A. Hernández-Segura, C. Jiménez-Mesa, Juan M. Górriz
AI summary
Summary: An Explainable Framework for the Relationship between Dementia and Metabolism PatternsOverview
Research area: Machine learning for neuroimaging — specifically, generative deep learning (variational autoencoders) applied to brain glucose metabolism imaging (FDG-PET) in Alzheimer's disease and dementia research.
Technical level: Intermediate. The paper assumes familiarity with variational autoencoders, latent variable models, and basic neuroimaging terminology, though its central idea — aligning one learned latent dimension with a clinical score — is explained in a largely self-contained way.
Scope (one sentence): The paper proposes a semi-supervised VAE that learns an interpretable latent biomarker of dementia severity from FDG-PET scans, disentangles it from confounding variability, and maps it back into brain space to reveal metabolic patterns.
Publication note: The manuscript states it has been peer-reviewed and published in NeuroImage (DOI: 10.1016/j.neuroimage.2026.121855), and that the official version should be cited instead of the preprint.
What This Paper Is About
High-dimensional neuroimaging data contains disease-relevant signal buried inside complex, non-linear variation, plus a great deal of nuisance variability from individual anatomy, scanner differences, and demographics. Clinical cognitive scores (such as ADAS13) measure impairment but do not say where in the brain the damage is; conventional image analysis can locate affected regions but does not produce a single continuous biomarker that tracks symptom severity.
This paper's goal is to bridge that gap: it builds a semi-supervised variational autoencoder that forces one latent variable (denoted z₀) to align with a dementia score while forcing another (z₁) to align with age, and then uses the model's generative ability to decode that biomarker back into brain space and visualize which regions lose metabolism as dementia progresses. Notably, the authors state they are not primarily interested in classification.
Key Contributions
-
A semi-supervised VAE framework with a flexible similarity regularization term. A third loss term encourages one or more chosen latent variables to be statistically similar to an external clinical or biomarker variable. The similarity metric (e.g., Pearson correlation, mutual information, Spearman correlation) and the target variable are both selectable, and the formulation is written generically as a similarity function D(z_(k), y).
-
An analysis of convergence behaviour identifying stable versus failure regimes. The authors construct phase diagrams over hyperparameters to show where the model collapses to the mean of the input data or to a trivial solution, and where it learns useful latent representations.
-
A generative mapping of neurodegenerative patterns back into brain space. Using average reconstructions decoded from the latent space plus a voxel-wise general linear model, the framework produces spatial maps of metabolic change associated with the dementia biomarker.
-
An interpretable account of the remaining, unsupervised latent variables. By systematically sampling and varying individual latent dimensions, the authors find the unsupervised dimensions encode affine transformations — rotation, translation, and scaling — plus intensity variations, which they interpret as confounding factors such as inter-subject variability and site-related noise.
Main Findings
-
z₀ tracks dementia severity. The first latent variable correlated with ADAS13 at |r| = 0.790 (p ≪ 0.001) and with the average FDG-PET score at |r| = 0.810 (p ≪ 0.001). Cognitive decline increased along z₀ while metabolism decreased along it.
-
The result is not tied to one similarity metric. Retraining with Spearman rank correlation instead of Pearson produced |ρ| = 0.77 (p ≪ 0.001) for ADAS13, which the authors describe as highly consistent with the Pearson result.
-
z₀ correlates with established structural biomarkers. Hippocampal volume |r| = 0.48, medial temporal lobes |r| = 0.45, entorhinal cortex |r| = 0.37, and fusiform volume |r| = 0.34 (all p ≪ 0.001). The text describes the hippocampus and medial temporal regions as showing robust associations whereas the fusiform gyrus showed only a weak correlation; the figure caption reports the 0.34 value. These structural measures were not part of the training objective and were used only for post hoc analysis.
-
Metabolic decline localizes to expected regions. The voxel-wise GLM map for z₀ showed decreased metabolism (blue) predominantly in the prefrontal and medial temporal cortices and parts of the occipital lobe, with a well-defined decline in the hippocampal region. Increased metabolism (red) appeared in the motor cortex and various subcortical structures. The authors note these structures were not explicitly given to the model and emerged post hoc.
-
Resting-state networks show differential involvement. The Default Mode Network and the left and right Fronto-Parietal Networks demonstrated significant reductions in metabolic activity, while the Sensorimotor Network exhibited either no significant changes or slight increases.
-
Classification performance is moderate and below the ADAS13 baseline. For AD vs. HC discrimination, the model achieved accuracy 0.8 ± 0.02, sensitivity 0.79 ± 0.04, specificity 0.77 ± 0.02, and balanced accuracy 0.79 ± 0.02 (bootstrap validation repeated 100 times). Using only ADAS13 as input gave accuracy 0.96 ± 0.03, sensitivity 0.95 ± 0.01, specificity 0.95 ± 0.02, and balanced accuracy 0.96 ± 0.02. Reported literature baselines: Wakefield et al. (2024) balanced accuracy 0.85 ± 0.01, and Dolci et al. (2024) accuracy 0.926 ± 0.02 with sensitivity 0.876 ± 0.03. The authors note these baselines were not re-implemented on their exact data split.
-
Two kinds of failure regime exist. Large β (strong KL regularization) or small latent dimensionality causes collapse toward the data mean. Large α for the similarity term makes the regularization overly dominant and yields non-informative representations; small α fails to produce dementia-correlated patterns. Both conditions must be satisfied for the framework to be considered effective.
-
Reconstructions are lower in intensity than inputs. The KL regularization counterweights the reconstruction, producing lower-quality scans than the input, though relevant structure is preserved.
Methodology in Plain English
Data. The study uses 3466 FDG-PET scans from the Alzheimer's Disease Neuroimaging Initiative (ADNI), measuring brain glucose metabolism, plus normalized structural biomarker volumes (hippocampus, medial temporal lobes, entorhinal cortex, fusiform). Because ADNI contains multiple longitudinal scans per person, splitting was done strictly at the subject level using unique participant identifiers to avoid data leakage. Subjects were randomly shuffled and assigned so that 65% of scans (n = 2239) went to training, 15% (n = 572) to validation, and 20% (n = 655) to test, with all scans from a given person staying in the same subset.
Preprocessing. PET scans were coregistered to a common template using rigid-body transformation (SPM12) but deliberately not spatially normalized, preserving individual anatomical variability for the convolutional architecture. Intensity normalization used min-max scaling relative to the 99th percentile followed by exponential transformations. Incomplete entries were excluded.
The model. A 3D convolutional variational autoencoder with a Beta-VAE style objective. The encoder has four convolutional layers with kernel sizes 11, 7, 5, and 3 (32, 64, 128, and 256 channels respectively), each followed by ReLU and batch normalization, feeding into a fully connected network that compresses 9216 neurons (256 channels × 3 × 4 × 3) into 256 neurons, then into separate mean and log-variance layers. The decoder maps the latent space back to 4608 neurons (128 channels × 3 × 4 × 3) and uses three transposed convolutional layers with kernel sizes 3, 4, and 11 (128, 64, and 32 channels).
The training objective. Three terms: mean squared error for reconstruction, a KL divergence term weighted by β, and the similarity term. The similarity term is the negative Pearson correlation between a chosen latent dimension and an external variable. In this work, the model was guided with −r(z₀, ADAS13) and −r(z₁, age). The authors emphasize the model is not trained to predict clinical scores directly — it is trained so that specific latent dimensions align with imaging patterns associated with those variables.
Chosen hyperparameters. Latent dimensionality 8, learning rate 2 × 10⁻⁵, batch size 8, β = 1 × 10⁻⁴, α_(j=0) = 2 × 10⁻⁴, α_(j=1) = 2 × 10⁻⁴, trained in PyTorch with Adam. The authors note similar values produce comparable results as long as they fall within the stable regimes.
Evaluation and interpretability. Three questions guide evaluation: does z₀ correlate with clinical and structural measures, do the patterns correspond to known AD-affected regions, and are confounds captured by other latent variables instead of z₀? The authors built phase diagrams to locate stable and failure regimes. To inspect brain-space effects, they generated synthetic subjects by varying one latent variable at a time while sampling the rest from a standard normal distribution N(0,1), decoded these into brain space, averaged them, and fitted a voxel-wise GLM. The GLM (equation 8) includes linear effects of each latent variable plus pairwise interaction terms; in this work it was restricted to L = 2, meaning only z₀ and z₁, because the remaining latents capture unsupervised variability not tied to predefined clinical factors. Regression coefficients were estimated with ordinary least squares and the β maps are visualized rather than thresholded statistical maps. Classification used binary logistic regression (HC = 0, AD = 1) with a maximum of 1000 iterations on latent representations, evaluated with accuracy, sensitivity, specificity, and balanced accuracy, and validated via bootstrap resampling repeated 100 times. An ablation study removed the similarity term entirely to compare guided versus unguided patterns.
Why This Matters
This work reframes what a generative neuroimaging model is for. Instead of chasing classification accuracy, it produces a single continuous, interpretable number that sits on a spectrum from healthy to impaired and can be decoded into a picture of where metabolism is failing. That combination — a scalar biomarker plus its spatial explanation — is what clinical scores and region-level analyses each provide only half of.
Real-world applications:
- Patient stratification and longitudinal tracking: A continuous latent biomarker that correlates with ADAS13 and structural volumes could be used to follow progression or select participants for trials, rather than relying on coarse diagnostic labels.
- Clinical trial enrichment: A biomarker shown to correlate with both cognitive and imaging measures could help define inclusion criteria or serve as a secondary endpoint.
- Interpretable reporting: Because the model's guidance variables are selectable, the same template could be re-targeted at other clinical measures or region volumes of interest.
- Multi-site imaging harmonization: The finding that unsupervised latent dimensions absorb affine transformations and intensity variations is directly relevant to studies where scanner and protocol differences contaminate the signal.
Industry relevance: The approach is relevant to pharmaceutical and imaging-biomarker companies that need interpretable, continuous readouts from PET data; to medical imaging software developers building explainable AI components; and to research consortia that need to distinguish biological signal from acquisition nuisance across sites.
Future Directions
-
Extending beyond two supervised dimensions. The GLM analysis was restricted to L = 2 (z₀ and z₁), leaving the remaining latent variables outside the regression. Whether other unsupervised dimensions carry clinically meaningful information that could be included in the analysis is left open.
-
Broadening the choice of similarity metric and target variable. The framework is formulated generically, and this work demonstrates only Pearson and Spearman correlation against ADAS13 and age. Applying it to other biomarkers, clinical scores, or region volumes — as the authors state is possible — remains to be tested systematically.
-
Improving discriminative performance. The model's balanced accuracy of 0.79 on AD vs. HC sits below the ADAS13-only baseline of 0.96 and below the reported literature baselines (0.85 and 0.926 accuracy figures). Whether the semi-supervised latent biomarker can be combined with other modalities to close this gap is unresolved.
-
Validating generalization outside ADNI. All experiments use ADNI FDG-PET data. The authors note that inter-study comparability with other FDG-PET ADNI-based methods is supported by their intensity normalization pipeline, but external validation on independent cohorts is not reported in this paper.
Target Audience
This paper is best suited to researchers and practitioners working at the intersection of deep learning and neuroimaging — particularly those interested in generative models for brain imaging, explainable AI in medicine, and biomarker discovery for neurodegenerative disease. It is also relevant to clinical researchers who want to understand what latent variable models can and cannot tell them about dementia progression, and to methodologists interested in how regularization hyperparameters shape whether a VAE learns something meaningful or collapses to a trivial solution. Readers without prior exposure to variational autoencoders will need to consult background material, but the paper's core logic — align one latent dimension to a clinical score, then decode it back to the brain — is accessible.
Authors’ abstract
High-dimensional neuroimaging data presents challenges for assessing neurodegenerative diseases due to complex non-linear relationships. Variational Autoencoders (VAEs) can encode scans into lower-dimensional latent spaces capturing disease-relevant features. We propose a semi-supervised VAE framework with a flexible similarity regularization term that aligns selected latent variables with clinical or biomarker measures of dementia progression. This allows adapting the similarity metric and supervised variables to specific goals or available data. We demonstrate the approach using PET scans from the Alzheimer's Disease Neuroimaging Initiative (ADNI), guiding the first latent dimension to align with a cognitive score. Using this supervised latent variable, we generate average reconstructions across levels of cognitive impairment. Voxel-wise GLM analysis reveals reduced metabolism in key regions, mainly the hippocampus, and within major Resting State Networks, particularly the Default Mode and Central Executive Networks. The remaining latent variables encode affine transformations and intensity variations, capturing confounds such as inter-subject variability and site effects. Our framework effectively extracts disease-related patterns aligned with established Alzheimer's biomarkers, offering an interpretable and adaptable tool for studying neurodegenerative progression.