Skip to content
AI.info

Research

GeMM-GAN: A Multimodal Generative Model Conditioned on Histopathology Images and Clinical Descriptions for Gene Expression Profile Generation

Overview Research area: Multimodal generative AI applied to computational pathology and transcriptomics (biomedical machine learning). Technical level: Advanced. The paper assumes familiarity with gen

GeMM-GAN: A Multimodal Generative Model Conditioned on Histopathology Images and Clinical Descriptions for Gene Expression Profile Generation
arXiv
2601.15392
Published
2026-01-21
Authors
Francesca Pia Panaccione, Carlo Sgaravatti, Pietro Pinoli

AI summary

Overview

  • Research area: Multimodal generative AI applied to computational pathology and transcriptomics (biomedical machine learning).
  • Technical level: Advanced. The paper assumes familiarity with generative adversarial networks, Wasserstein GANs with gradient penalty, Transformer encoders, cross-attention, Feature-wise Linear Modulation (FiLM), and histopathology whole-slide image preprocessing.
  • Scope: The paper introduces and empirically evaluates GeMM-GAN, a generative adversarial framework that synthesizes gene expression profiles conditioned jointly on histopathology tissue slides and clinical text descriptions, tested on a 1,944-case subset of the TCGA dataset covering 19 tumor types and 18,868 genes.

What This Paper Is About

Gene expression profiling is valuable for research and treatment selection, but it is limited by cost, privacy regulations, and uneven availability across medical centers, while histopathology slides and clinical metadata are collected routinely. The paper's goal is to close that data gap by training a generative model that takes a whole-slide histopathology image plus a written clinical description and outputs a biologically plausible gene expression profile, so that synthetic transcriptomic data can be produced without performing RNA sequencing.

Key Contributions

  1. A novel multimodal generative framework. The authors present GeMM-GAN for generating gene expression profiles conditioned jointly on histopathology images and clinical descriptions, which they state is the first approach to condition a generative model of gene expression on both modalities simultaneously.
  2. Exploration of cross-modal fusion strategies. The paper investigates integrating text into the image-based pipeline through Feature-wise Linear Modulation (FiLM) of patch embeddings and a bidirectional cross-attention mechanism between text and patch tokens, producing a shared latent conditioning vector.
  3. End-to-end multimodal conditioning of a WGAN-GP. The fused multimodal embedding conditions a Wasserstein GAN with Gradient Penalty, with separate Multimodal Fusion networks and learnable parameters (theta_G for the generator, theta_D for the discriminator).
  4. A full empirical evaluation and ablation. The method is compared against a Vanilla WGAN-GP, a Conditional WGAN-GP, and a Conditional VAE across unsupervised metrics, detectability, and downstream utility, plus an ablation isolating the contributions of text, images, FiLM, and cross-attention.

Main Findings

  • Recall and distributional coverage: GeMM-GAN achieves a recall of 0.8190, higher than the Vanilla WGAN-GP (0.6477), the Conditional WGAN-GP (0.7372), and the CVAE (0.7956), which the authors interpret as capturing dataset variability and avoiding mode collapse.
  • Precision trade-off: Precision is 0.7103, which is lower than the Conditional WGAN-GP (0.8436) and the Vanilla WGAN-GP (0.8282), but higher than the CVAE (0.4892).
  • Gene correlation preservation: The proposed method reaches a correlation MSE (C. MSE) of 0.0007, which the paper describes as a 98.2% improvement over the Conditional WGAN-GP value of 0.0038. Other baselines scored 0.0095 (Vanilla WGAN-GP) and 0.2837 (CVAE).
  • Detectability by logistic regression: Synthetic samples were much harder to distinguish from real data under logistic regression, with accuracy 0.5236 and F1 0.5444, versus 0.8459/0.8662 (Vanilla WGAN-GP), 0.8801/0.8925 (Conditional WGAN-GP), and 0.8931/0.9003 (CVAE).
  • Detectability by MLP remains high: A multilayer perceptron still distinguishes generated from real samples with accuracy 0.9797 and F1 0.9794, which the authors flag as an area for future investigation.
  • Downstream disease type utility: For disease type classification, the method achieves 16.9% improvement in Random Forest accuracy and 11.1% improvement in F1 score relative to the baselines, reaching RF accuracy 0.7449 and RF F1 0.8249 and MLP accuracy 0.9354 and MLP F1 0.9322.
  • Downstream primary site utility: Primary site classification shows 13.4% improvement in RF accuracy and 10.0% improvement in F1 score, reaching RF accuracy 0.6936, RF F1 0.7478, MLP accuracy 0.8436, and MLP F1 0.8264.
  • Ablation: text-only conditioning is selective but narrow: Conditioning on the text CLS token alone gives high precision (0.9513) but markedly lower recall (0.5205), and a utility RF accuracy of 0.5000 and F1 of 0.6330.
  • Ablation: image conditioning drives utility: Image-informed variants produce more stable precision and recall values and consistently higher accuracy and F1 on downstream tasks; the full model (precision 0.7103, recall 0.8190, C. MSE 0.0007) achieves the best overall utility (RF accuracy 0.7449, F1 0.8249) and detectability (LR accuracy 0.5236, F1 0.5444).
  • Abstract-level claim: The framework improves accuracy on downstream disease type prediction by more than 11% compared with current state-of-the-art generative models.

Methodology in Plain English

The pipeline has three stages, all trained jointly end to end.

  1. Extract and encode single modalities. Tissue slides are segmented with Otsu thresholding, cut into 256 x 256 tiles, and only tiles with more than 20% tissue content are kept. At each training step, 256 patches are randomly sampled and passed through UNI, a pretrained Vision Transformer for histopathology tiles. Clinical metadata (JSON with demographics, cancer subtype, treatment history) is converted into roughly 200-word case summaries by a version of Llama3-8B fine-tuned on medical data, with irrelevant fields removed, and those summaries are encoded by Clinical ModernBERT. Because UNI outputs 1024-dimensional embeddings and Clinical ModernBERT outputs 768-dimensional ones, linear projection layers align both to d = 256; the pretrained backbones are frozen and only the projections are trained.
  2. Fuse the modalities. A learnable patch CLS token is added to the image side since patch encoding produces no natural CLS token. FiLM uses the text CLS token to apply a feature-wise affine transformation (a learned scale gamma and shift beta) to the patch embeddings, emphasizing features relevant to the clinical description. The modulated patches plus the patch CLS token go through a Transformer Encoder to model relationships across patches. Then bidirectional cross-attention runs in both directions: the text CLS token attends over the visual tokens (Text2Image), and the updated patch CLS token attends over the text tokens (Image2Text). The two resulting vectors are summed into one multimodal embedding.
  3. Generate the expression profile. A WGAN-GP is conditioned on this embedding. Two separate Multimodal Fusion networks with different learnable parameters produce conditioning vectors for the generator and the discriminator; both are multilayer perceptrons with two hidden layers of size 256. The generator takes a noise vector sampled from a normal distribution plus the conditioning vector and outputs a gene expression profile.

Data and setup. Everything comes from TCGA, restricted to samples whose tissue slide is under 100 MB, giving 1,944 clinical cases across 19 tumor types with expression values for 18,868 genes, quantified as FPKM and z-score normalized after removing genes with more than 90% missing values. An 80/20 train-test split keeps images, metadata, and expression profiles aligned. Training used an AMD Ryzen 1950X CPU with 128 GB RAM and two NVIDIA A6000 GPUs with 48 GB VRAM, PyTorch 2.6.0 and CUDA 12.4, a latent dimension of 256, and a batch size of 64.

Evaluation. Three metric families were used: unsupervised metrics (precision and recall computed with the 10th nearest neighbor, plus MSE between real and synthetic gene-gene correlation matrices), detectability (logistic regression and MLP classifiers trying to separate real from generated samples; lower is better), and utility (classifiers trained on generated data and tested on real data to predict disease type or primary site). Baselines were a Vanilla WGAN-GP, a Conditional WGAN-GP conditioned on disease type and primary site, and a Conditional VAE conditioned on the same two variables; utility could not be computed for the Vanilla WGAN-GP because it is unconditional. Results are reported as means and standard deviations over 10 generation runs.

Why This Matters

The work targets a practical asymmetry in biomedical data: images and clinical notes are abundant, transcriptomic profiles are not. If expression profiles can be synthesized plausibly from what hospitals already collect, downstream research and model development become possible without new sequencing costs or privacy exposure. The paper claims to be the first to condition transcriptomic generation on images and text together, moving beyond earlier work that used only low-dimensional vectors, only pathology images, or framed the task as pure prediction rather than generation.

Real-world applications:

  • Data augmentation for rare tumor subtypes, where sequencing data are too scarce to train reliable models.
  • Trial enrichment and cohort simulation, using synthetic profiles to explore disease type and primary site distributions before committing to sequencing.
  • Clinical decision support around expression-based assays, such as PAM50-guided treatment decisions in breast cancer, which the paper cites as an example where expression analysis improves survival outcomes through risk stratification.
  • Privacy-preserving data sharing, since synthetic profiles can be distributed more freely than identifiable patient transcriptomes.

Industry relevance: Pharmaceutical and diagnostics companies, hospital systems with digital pathology archives, and computational pathology vendors all face the same data gap. A model that converts routinely collected slides and notes into usable molecular representations could reduce reliance on costly RNA sequencing, accelerate biomarker discovery, and support cross-modal AI tools for precision medicine. The authors state that code will be available at a public GitHub repository.

Future Directions

  • Closing the detectability gap. An MLP still separates real from generated profiles with accuracy 0.9797 and F1 0.9794, while logistic regression is near chance (0.5236 accuracy). The authors explicitly identify sophisticated classifiers as an area for future investigation.
  • Extending the framework to image generation. Future work aims to generate histopathology images from gene expression profiles and clinical data, creating a bidirectional capability that would let researchers simulate how genetic changes appear in tissue.
  • Refining the precision-recall balance. The method trades precision for recall relative to the Conditional WGAN-GP; the ablation shows text-only conditioning pushes precision to 0.9513 at the cost of recall (0.5205), so how to combine both without sacrificing either remains open.
  • Scaling beyond the current data subset. Evaluation was restricted to slides under 100 MB, 1,944 cases, 19 tumor types, and 18,868 genes; whether the approach holds on larger and more heterogeneous cohorts is not established here.

Target Audience

This paper is most useful to researchers and practitioners working at the intersection of generative AI and biomedical data: machine learning engineers building synthetic data pipelines for omics, computational pathologists interested in multimodal modeling of whole-slide images, and bioinformaticians evaluating whether synthetic transcriptomes are usable for downstream tasks. It also suits clinical informatics groups and industry teams exploring alternatives to costly RNA sequencing, and graduate-level readers with prior exposure to GANs and Transformer architectures who want a concrete case study in multimodal biomedical generation.

Authors’ abstract

Biomedical research increasingly relies on integrating diverse data modalities, including gene expression profiles, medical images, and clinical metadata. While medical images and clinical metadata are routinely collected in clinical practice, gene expression data presents unique challenges for widespread research use, mainly due to stringent privacy regulations and costly laboratory experiments. To address these limitations, we present GeMM-GAN, a novel Generative Adversarial Network conditioned on histopathology tissue slides and clinical metadata, designed to synthesize realistic gene expression profiles. GeMM-GAN combines a Transformer Encoder for image patches with a final Cross Attention mechanism between patches and text tokens, producing a conditioning vector to guide a generative model in generating biologically coherent gene expression profiles. We evaluate our approach on the TCGA dataset and demonstrate that our framework outperforms standard generative models and generates more realistic and functionally meaningful gene expression profiles, improving by more than 11\% the accuracy on downstream disease type prediction compared to current state-of-the-art generative models. Code will be available at: https://github.com/francescapia/GeMM-GAN

Read the original paper