Research
Variational decomposition autoencoding improves disentanglement of latent representations
Overview Research area: Unsupervised representation learning and disentanglement for nonstationary, high-dimensional time-evolving signals, with applications in speech processing and biomedical signal

- arXiv
- 2601.06844
- Published
- 2026-01-11
- Authors
- Ioannis Ziogas, Aamna Al Shehhi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis
AI summary
Overview
Research area: Unsupervised representation learning and disentanglement for nonstationary, high-dimensional time-evolving signals, with applications in speech processing and biomedical signal analysis.
Technical level: Advanced. The paper assumes familiarity with variational autoencoders, evidence lower bounds, variational inference, contrastive self-supervised learning, and classical signal decomposition methods.
Scope: The paper introduces variational decomposition autoencoding (VDA) and its encoder-only neural implementation, the variational decomposition autoencoder (DecVAE), which embeds a signal-decomposition structural bias into the variational generative process so that latent subspaces align with time-frequency structure.
What This Paper Is About
Standard variational autoencoders assume that the generative factors of data are statistically independent and can be pushed into separate latent dimensions by approximating a multivariate Gaussian prior. That assumption breaks down for complex, nonstationary signals such as speech, where factors like speaker identity, phonetic content, and clinical state are numerous, interdependent, and not cleanly defined. The paper's goal is to build a representation learning framework in which decomposition itself—splitting the input into frequency-localized components—is baked into the generative process and the training objective, producing latent subspaces that are disentangled, interpretable, and transferable to new tasks and domains.
Key Contributions
-
The VDA framework. A reformulation of the VAE generative model in which the latent prior is treated as a composition of multiple frequency-related priors, and a single recognition model approximates those subspaces conditionally using information from the other subspaces and the observed input.
-
The DecVAE architecture. An encoder-only network (implemented with one-dimensional convolution layers) that combines a signal decomposition model, a self-supervised contrastive task that drives latent components toward orthogonality, and variational prior approximation within each subspace. Notably, the decoder is deliberately excluded to isolate the effect of structural biases.
-
The decomposition evidence lower bound (DELBO). A training objective that extends the classic evidence lower bound with a self-supervised loss embodying the decomposition structural bias, including orthogonality and component-level reconstruction mechanics.
-
Empirical validation across four datasets plus a theoretical account. Evaluation on SimVowels, TIMIT, VOC-ALS, and IEMOCAP spans zero-shot transfer and fine-tuning, with ablations over decomposition method, number of components, and the β parameter. The authors also argue that VDA is a hybrid of variational and PCA-based methods and can be formulated as a singular value decomposition operation on the input. All R/Python code is stated to be publicly available.
Main Findings
-
Simulated speech (SimVowels): With a decomposition of C=3 components, DecVAE disentangled the vowel and speaker factors in both frame-level and sequence-level latent spaces better than a frequency-insensitive VAE, producing smoother, better-separated embeddings and a more elaborated latent structure. DecVAE representations scored higher on DCI and modularity-explicitness and showed higher IRS robustness than VAE-based models, ICA, and PCA. Increasing to C=4 gave only minor improvements.
-
Decomposition choice matters: Comparing the authors' custom filter decomposition (FD) against empirical mode decomposition (EMD), variational mode decomposition (VMD), and empirical wavelet transform (EWT), FD and EWT performed best because filters and wavelet transforms promote time-frequency orthogonality. With VMD and EMD, correlations in the input propagated into the latent space and contaminated factor separation, worsening both disentanglement and task metrics.
-
Effect of β: DecVAE performance generally declined as β increased, with better results in the low range (0.1, 1) and an optimal β=0.1. DecVAE models were more robust to β than β-VAE models because of the additional disentanglement dynamics, and DecVAEs could still disentangle when β=0.
-
Real speech (TIMIT): Using C=4, DecVAE extended frequency-resonant embedding to real speech containing consonants beyond vowels, learning smoother, more Gaussian latent spaces where both phonemes and speakers were separated. DecVAEs showed higher disentanglement, informativeness, robustness, and explicitness but reduced modularity relative to other methods. β-VAE required fewer dimensions (higher completeness), and PCA learned the most uncorrelated latent space without disentangling generative factors. Raw Mel filterbank inputs were already quite informative for speech recognition. Increasing C helped all disentanglement metrics.
-
Zero-shot dysarthria severity (VOC-ALS): A β-DecVAE pretrained on SimVowels (β=0.1) transferred without adaptation and retained a frequency-resonant latent structure on unseen data, disentangling three generative factors—phoneme, speaker identity, and ALS progression as reflected in King's clinical stage. VAE-based and other methods failed to separate all three. DecVAEs outperformed other methods in all disentanglement and task-specific metrics except completeness, with PCA and ICA learning the most uncorrelated latents by MI and GCN. King's clinical stage was treated as a fixed per-individual reference point, so the analysis focused on how phonetic content varies across stages rather than intra-speaker change over time.
-
Emotional speech (IEMOCAP): Because emotion has no direct time-frequency correspondence, the input and its decomposition offered no separation for emotion or phonetic content. After transfer and fine-tuning, DecVAE adapted to the new distribution and learned a frequency-resonant embedding with enhanced emotion, speaker, and phoneme disentanglement relative to β-VAE. DecVAEs outperformed other methods on generalization/classification and on disentanglement, informativeness, and robustness, but not on modularity, MI, or GCN, where PCA and ICA produced the most uncorrelated latent spaces. The sequence branch contributed to emotion detection performance.
-
Consistent weakness in compactness: DecVAE consistently underperformed on compactness as measured by DCI completeness. In VAEs, compactness arises from the combination of Gaussian prior approximation and decoder reconstruction; the paper observes that this yields compact yet entangled and uninformative representations, whereas DecVAE representations are not compact but are disentangled and informative.
-
Ablations on the Gaussian prior: Disentanglement arose even without Gaussian prior approximation, because DecVAE consistently maximizes and minimizes divergences of components against other components and the original signal in both real and simulated data.
-
Training sensitivities: Convergence slows when the number of components C is large, when very few frames are used in the self-supervised loss, and when the input is more abstract, such as raw waveform instead of Mel filterbanks. These conditions harm both disentanglement and informativeness.
-
Interpretability: Latent traversal analysis showed DecVAE splitting information across latent subspaces of oscillatory components, giving complementary allocation across generative factor values, whereas VAE and β-VAE models compress information into a few dimensions or a single latent space, limiting separability.
-
Reporting note: The provided text states that DecVAEs surpass state-of-the-art methods across tasks and reports relative metric comparisons and error-bar conventions (95% confidence intervals in Figures 4h and 5h), but does not include the specific numerical accuracy, F1, or disentanglement score values.
Methodology in Plain English
The researchers begin from the observation that a VAE's push toward an uncorrelated Gaussian latent space is a weak structural bias for signals whose factors overlap in time and frequency. Their alternative treats the latent prior as built from several frequency-related pieces rather than one monolithic Gaussian. An input time series is first passed through a signal decomposition model that splits it into C components, each capturing localized time-frequency content. Each component is then encoded into its own latent subspace, so the latent space is structured rather than flat.
Training uses a self-supervised contrastive objective on top of the usual variational objective. This objective does two things: it encourages the latent subspaces to be orthogonal to each other while still reconstructing their own components accurately, and it regularizes each subspace toward a Gaussian prior. Because the architecture is encoder-only, the authors rely on these biases rather than on decoder reconstruction to produce disentanglement.
The model can also operate at two time scales at once—a long-term sequence branch and a short-term frame branch—with separate encoders and aggregation functions combining the subspaces, mirroring the idea that slow variables like speaker identity shape fast variables like phonetic content.
Evaluation compares DecVAE against VAE, β-VAE, ICA, and PCA, plus other disentanglement methods discussed in related work, using multiple disentanglement metrics (DCI, modularity-explicitness, IRS robustness, MI, GCN) alongside task metrics for phoneme recognition, speaker identification, disease-stage prediction, disease duration, and emotion classification. Transfer is tested in two modes: zero-shot (pretrained on SimVowels, applied unchanged) and fine-tuned.
Why This Matters
Impact on research. The paper argues that structural biases—specifically, encoding decomposition into the generative process—are what mostly promote disentanglement, rather than the Gaussian prior and decoder alone. It supports this by removing the decoder entirely and still achieving state-of-the-art disentanglement, and by positioning VDA as a variational analogue of PCA-style orthogonal component pursuit. This reframes disentanglement as a property that can be engineered through architectural and objective design rather than only through prior regularization.
Real-world applications (as stated or directly implied):
- Clinical diagnostics, including evaluation of dysarthria severity and disease progression from speech.
- Human-computer interaction, including speech recognition and emotional speech classification.
- Adaptive neurotechnologies.
- Speech recognition for both healthy and dysarthric speakers, and speaker identification.
Industry relevance. The model is a lightweight convolutional encoder with no decoder, trained self-supervised, that can be pretrained once and then either used zero-shot or fine-tuned on new domains with simple downstream classifiers. That profile suits deployment in low-data clinical or affective computing settings, and the authors note the design is modality-agnostic and applicable to any sequence modality, opening a path to multi-modal fusion where each modality is treated as its own component.
Future Directions
- Generation and decoder-based modeling. The authors excluded the decoder to demonstrate the strength of the structural biases and state that they focused evaluation on disentanglement quality rather than generation; high-quality generation is left as future work.
- Multi-modal disentanglement. Using modalities (audio, visual, textual, physiological) as components in place of a signal decomposition, so DecVAE learns modality-specific latent subspaces while preserving cross-modal coherence, and supports fusion tasks.
- Optimization and scaling. Convergence slows with a large number of components, with too few frames in the self-supervised loss, and with abstract inputs such as raw waveform; resolving these bottlenecks is an open problem.
- Broader scientific domains and interpretability. Extending the framework beyond speech to other nonstationary time series and further developing latent response analysis to reveal how individual subspaces encode specific generative factors.
Target Audience
Researchers and practitioners in machine learning and signal processing who work on disentangled representation learning, variational inference, or self-supervised learning for time series. It is also relevant to speech scientists and biomedical engineers interested in interpretable embeddings for applications such as dysarthria assessment and emotional speech analysis, and to engineers who need pretrained, transferable encoders for sequence data. Readers without a background in variational methods or signal decomposition will find the technical sections difficult.
Authors’ abstract
Understanding the structure of complex, nonstationary, high-dimensional time-evolving signals is a central challenge in scientific data analysis. In many domains, such as speech and biomedical signal processing, the ability to learn disentangled and interpretable representations is critical for uncovering latent generative mechanisms. Traditional approaches to unsupervised representation learning, including variational autoencoders (VAEs), often struggle to capture the temporal and spectral diversity inherent in such data. Here we introduce variational decomposition autoencoding (VDA), a framework that extends VAEs by incorporating a strong structural bias toward signal decomposition. VDA is instantiated through variational decomposition autoencoders (DecVAEs), i.e., encoder-only neural networks that combine a signal decomposition model, a contrastive self-supervised task, and variational prior approximation to learn multiple latent subspaces aligned with time-frequency characteristics. We demonstrate the effectiveness of DecVAEs on simulated data and three publicly available scientific datasets, spanning speech recognition, dysarthria severity evaluation, and emotional speech classification. Our results demonstrate that DecVAEs surpass state-of-the-art VAE-based methods in terms of disentanglement quality, generalization across tasks, and the interpretability of latent encodings. These findings suggest that decomposition-aware architectures can serve as robust tools for extracting structured representations from dynamic signals, with potential applications in clinical diagnostics, human-computer interaction, and adaptive neurotechnologies.