Skip to content
AI.info

Research

XFACTORS: Disentangled Information Bottleneck via Contrastive Supervision

XFactors: Disentangled Information Bottleneck via Contrastive Supervision Overview Research area: Representation learning / disentangled generative modeling, spanning unsupervised and weakly supervise

XFACTORS: Disentangled Information Bottleneck via Contrastive Supervision
arXiv
2601.21688
Published
2026-01-29
Authors
Alexandre Myara, Nicolas Bourriez, Thomas Boyer, Thomas Lemercier, Ihab Bendidi, Auguste Genovesio

AI summary

XFactors: Disentangled Information Bottleneck via Contrastive Supervision

Overview

Research area: Representation learning / disentangled generative modeling, spanning unsupervised and weakly supervised VAEs, information-theoretic objectives, and contrastive learning (with applications to vision benchmarks and biological imaging).

Technical level: Intermediate (the method's mechanics are accessible with a working knowledge of VAEs and contrastive learning; the information-theoretic framing of the objective in Section 3.1 is more advanced).

Scope: The paper introduces a weakly supervised VAE that isolates a chosen subset of labeled factors into dedicated latent blocks while preserving all remaining variation in a residual block, and evaluates it on four synthetic disentanglement benchmarks plus CelebA and JUMP Cell Painting.

What This Paper Is About

Deep networks often learn features that predict meaningful attributes, but those attributes are typically spread across many dimensions rather than exposed as stable, addressable variables that a user can inspect or replace. At the same time, fully disentangling every latent cause is unrealistic: in most real datasets only a few factors are labeled, relevant, or actionable, while the rest are unknown, irrelevant, or too numerous to isolate. XFactors addresses this by letting a user specify in advance which labeled factors should become controllable, assigning each to its own latent block and letting a KL-regularized residual space absorb everything else.

Key Contributions

  1. Targeted partial disentanglement formulation. The paper defines a setting in which a selected subset of labeled factors is made controllable while non-targeted variation is fully retained in a residual subspace, rather than being discarded or forced into the supervised blocks.

  2. The XFactors architecture. A weakly supervised VAE whose latent space is decomposed as a direct sum of factor-specific blocks ($\mathcal{T}_1, \ldots, \mathcal{T}_K$) and a residual subspace ($\mathcal{S}$), supervised per block by InfoNCE contrastive losses and regularized by reconstruction and KL terms — with no adversarial training and no auxiliary classifier serving as the main disentanglement mechanism.

  3. An adapted objective (X-DisenIB). A reformulation of the Disentangled Information Bottleneck (DisenIB, Pan et al., 2020) that sums one mutual-information term per target factor, $\mathcal{I}(T_i; y_{f_i})$, replacing adversarial min-max games with a direct minimization problem that scales to multiple target factors by assigning one contrastive signal per block.

  4. Empirical validation across domains. Strong disentanglement scores on 3DShapes, MPI3D, dSprites, and Cars3D with a single fixed hyperparameter set, factor-swapping generations, ablation and scaling studies, and qualitative proof-of-concept results on CelebA and JUMP Cell Painting.

Main Findings

  • State-of-the-art-level scores on fully disentangled benchmarks. On Shapes3D, XFactors reports FactorVAE 1.000 ± .000 with $|S|=1$ and 1.000 with $|S|=3$, and DCI 1.000 ± .000 and 1.000 respectively. On MPI3D it reports FactorVAE .978 ± .001 ($|S|=1$) and .999 ($|S|=3$), with DCI .949 ± .000 and .980. On dSprites it reports FactorVAE .952 ± .003 and .990, with DCI .909 ± .001 and .824. On cars3D it reports FactorVAE .948 ± .002 and 1.000, with DCI .802 ± .004 and .885. The authors note that on these datasets, where direct comparison is meaningful, XFactors obtains the highest reported scores in the comparisons.

  • One hyperparameter set across all datasets. Except for the scaling and ablation studies, a single configuration is used everywhere: $\lambda_{\mathrm{NCE}} = 0.5$, $\dim \mathcal{T}_i = 2$, $\dim \mathcal{S} = 126$, $\beta_s = 100$, $\beta_t = 100$ — no dataset-specific tuning.

  • CelebA results require careful reading. XFactors supervises 5 attributes and leaves the rest of the facial variation in $\mathcal{S}$, reporting FactorVAE .532 ± .002 and DCI .690 ± .001. The paper explicitly frames this table as not a direct full-disentanglement benchmark, since methods differ in supervision and target different numbers of attributes (CMI disentangles only 2 attributes; the unsupervised baselines operate on 40).

  • Factor swapping works as a direct intervention. Replacing one $\mathcal{T}_i$ block in a source representation with the corresponding block from a target image often changes the intended attribute while largely preserving the rest of the representation — shown qualitatively on CelebA, MPI3D, dSprites, and Shapes3D. The authors emphasize these are not photorealistic syntheses and that the architecture remains a simple VAE prioritizing disentanglement over reconstruction quality.

  • Residual capacity does not degrade the interface when targets are well regularized. Increasing $\dim S$ up to 126 does not reduce the FactorVAE score when the target subspace is adequately regularized. Under weak target regularization ($\beta_t = 1$) performance drops at higher dimensions, while $\beta_t = 100$ mitigates this decay.

  • Ablations confirm both components matter. On Shapes3D, removing InfoNCE drops FactorVAE from 1.000 ± .000 to .964 ± .046 and D from 1.000 ± .000 to .933 ± .078; removing $\mathcal{S}$ drops FactorVAE to .833 ± .000 and D to .829 ± .005. Increasing $\dim \mathcal{T}$ from 2 to 3 does not meaningfully change performance (FactorVAE 1.000 ± .000, D 1.000 ± .000).

  • Qualitative biological proof of concept. On JUMP Cell Painting, compound identity and experimental source are assigned to two distinct $\mathcal{T}$ subspaces for 8 positive-control compounds and DMSO controls across 7 experimental sources, with remaining biological and technical variation left in $\mathcal{S}$. The paper describes this as a qualitative stress test rather than a complete biological validation.

  • Latent visualization shows structure. On MPI3D, each $\mathcal{T}_i$ subspace (2D) shows clear structure when colored by its target factor, and the freely varying factor is also represented in $\mathcal{S}$ without explicit supervision, consistent with the organizing effect of KL regularization. The $\mathcal{S}$ plot is described as less contrasted because it is a projection with 67% variance explained.

  • Stated limitation: reconstruction quality. The main limitation acknowledged is reconstruction quality, especially on CelebA, attributed partly to the classical VAE trade-off between reconstruction fidelity and KL regularization — though the authors state it does not affect the disentanglement objective addressed.

Methodology in Plain English

The approach starts from an information-theoretic view of representation learning: an encoder should compress the input while keeping task-relevant information. Prior work used this to split a latent space into a "target" part and a "nuisance" part, but typically enforced the split with adversarial min-max games, which are unstable and hard to scale.

XFactors keeps the structural split but changes how it is enforced. The latent space is written as a direct sum of $K$ small factor blocks $\mathcal{T}_1, \ldots, \mathcal{T}_K$ plus a residual subspace $\mathcal{S}$. Two parallel encoders handle the two roles: $\psi_s(\cdot)$ produces the residual code $\boldsymbol{z}_s$ and $\psi_t(\cdot)$ produces the factor code $\boldsymbol{z}_t$. The two codes are concatenated and passed to a single decoder $\phi(\cdot)$.

Each factor block is trained with a supervised InfoNCE loss: within a batch, the code of an anchor in block $\mathcal{T}i$ is pulled toward codes sharing the same label value for that factor and pushed away from the rest, using a temperature $\tau$. Anchors with no positive in the batch are ignored for that factor. The paper notes this yields the motivational relation $I(T_i; y{f_i}) \geq \log N - \mathcal{L}_{\mathrm{InfoNCE}}$ with $N = |T_i|$, but is used as a practical contrastive surrogate rather than exact mutual-information maximization.

The implemented loss is a sum of four terms: an $L^2$ reconstruction loss, a KL term for $\mathcal{S}$ weighted by $\beta_s$, a KL term for the aggregated factor subspace weighted by $\beta_t$, and the sum of per-factor InfoNCE losses weighted by $\lambda_i$. Both KL terms push the branches toward Gaussian priors, which the authors describe as organizing pressures supporting interpolation and sampling rather than certificates of statistical independence.

A key design choice is the residual block. If all latent capacity went to the target factors, unannotated variation would either be lost — degrading reconstructions and downstream use — or leak into the supervised blocks. Giving non-targeted variation a regularized destination means $\dim(\mathcal{S})$ can be large enough to absorb nuisance or content variation without contaminating the factor blocks. The separation between $\mathcal{S}$ and $\mathcal{T}$ is encouraged by the latent partition, the contrastive block losses, the reconstruction bottleneck, and KL regularization, not by an explicit estimator of $\mathcal{I}(T;S)$.

At inference, controlled editing is simple: encode a source and a target image, keep the source's residual code $\boldsymbol{z}_s$, substitute the target's code for the chosen factor block, keep the other factor blocks from the source, and decode.

Why This Matters

Impact on research. The paper argues that the field's implicit goal of recovering all latent causes is rarely realistic, and that existing benchmarks blur the distinction between fully factorized synthetic datasets with complete factor coverage and realistic datasets like CelebA where attributes are correlated and coverage is incomplete. It offers an alternative in which supervision is applied only where it is available and actionable, and it deliberately avoids two known pitfalls: classifier-based supervision, whose decision boundaries are sensitive to adversarial perturbations, and unsupervised methods that require a tedious post-hoc search over latent directions before each coordinate can be linked to a semantic factor. Replacing adversarial games with a direct minimization objective also simplifies training and lets the method scale to multiple target factors.

Real-world applications.

  • Attribute editing with a controlled handle: replacing a named latent block edits a chosen attribute while holding the remaining latent components fixed.
  • Screening and perturbation biology: the JUMP Cell Painting experiment targets compound perturbation and experimental source in separate blocks for 8 positive-control compounds and DMSO controls across 7 experimental sources — useful for source or batch-effect inspection and removal.
  • Nuisance-factor analysis and dataset debugging: the residual space gives uncontrolled variation a place to live and can be inspected.
  • Factor-specific retrieval and counterfactual generation: the named blocks make retrieval and controlled counterfactual queries possible without searching the latent space.

Industry relevance. Any pipeline where a small set of metadata fields is reliably labeled — drug compound, cell line, acquisition site, product category, demographic attribute — but everything else is unannotated can use the same pattern: expose the fields that matter as addressable variables, and let a regularized residual absorb the rest so downstream models and humans can intervene without retraining a bespoke head. The fact that a single hyperparameter configuration holds across four synthetic datasets plus CelebA and JUMP Cell Painting lowers the cost of adopting it in a new domain.

Future Directions

  1. Stronger generative backbones. Because the main stated limitation is reconstruction quality — especially on CelebA — the paper suggests combining the factor interface with more powerful generative backbones rather than a simple VAE.

  2. More systematic evaluation of partial disentanglement. The authors call for broader leakage diagnostics between $\mathcal{S}$ and the $\mathcal{T}_i$ blocks, correlation-shift tests, and held-out-combination evaluations, since existing benchmarks blend fully factorized datasets with realistic, incompletely covered ones.

  3. Explicit measurement of the subspace separation. In the current implementation, $\mathcal{I}(T;S)$ is only an idealized motivation — separation is encouraged by the partition, contrastive losses, bottleneck, and KL terms rather than estimated directly. Estimating or bounding it is a natural next step.

  4. Biological applications beyond proof of concept. The JUMP results are explicitly framed as qualitative; moving toward source or batch-effect removal and trustworthy counterfactual generation in real biological workflows is left open.

Target Audience

Researchers and practitioners who need controllable, inspectable representations rather than purely predictive ones: disentangled and causal representation learning researchers, generative modeling groups working with VAEs and contrastive objectives, and applied scientists in computational biology or any domain with partial, incomplete annotation of factors — such as screening datasets that include drug compounds across a limited set of cell lines where exhaustive combinations are prohibitively expensive. Readers who want a ready-to-use method will need moderate familiarity with VAEs, KL regularization, and InfoNCE; readers interested in the information-theoretic justification of the objective should be comfortable with mutual information notation.

Authors’ abstract

Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, purely unsupervised approaches have proven successful on fully disentangled synthetic data, but fail to recover semantic factors from real data without strong inductive biases. On the other hand, supervised approaches are unstable and hard to scale to large attribute sets because they rely on adversarial objectives or auxiliary classifiers. We introduce \textsc{XFactors}, a weakly-supervised VAE framework that disentangles and provides explicit control over a chosen set of factors. Building on the Disentangled Information Bottleneck perspective, we decompose the representation into a residual subspace $\mathcal{S}$ and factor-specific subspaces $\mathcal{T}_1,\ldots,\mathcal{T}_K$ and a residual subspace $\mathcal{S}$. Each target factor is encoded in its assigned $\mathcal{T}_i$ through contrastive supervision: an InfoNCE loss pulls together latents sharing the same factor value and pushes apart mismatched pairs. In parallel, KL regularization imposes a Gaussian structure on both $\mathcal{S}$ and the aggregated factor subspaces, organizing the geometry without additional supervision for non-targeted factors and avoiding adversarial training and classifiers. Across multiple datasets, with constant hyperparameters, \textsc{XFactors} achieves state-of-the-art disentanglement scores and yields consistent qualitative factor alignment in the corresponding subspaces, enabling controlled factor swapping via latent replacement. We further demonstrate that our method scales correctly with increasing latent capacity and evaluate it on the real-world dataset CelebA. Our code is available at \href{https://github.com/ICML26-anon/XFactors}{github.com/ICML26-anon/XFactors}.

Read the original paper