Skip to content
AI.info

Research

Conditional Neural Optimal Transport for Predicting Cellular Phenotypes from Molecular Structure

Overview Research area: Computational biology / computer vision applied to high-content microscopy — specifically, predicting image-derived cellular phenotypes from chemical structure. Technical level

arXiv
2608.14293
Published
2026-08-14
Authors
Gauthier Avité, Maxime Sanchez-Renauld, Nicolas Bourriez, Auguste Genovesio

AI summary

Overview

Research area: Computational biology / computer vision applied to high-content microscopy — specifically, predicting image-derived cellular phenotypes from chemical structure.

Technical level: Advanced. The paper assumes familiarity with optimal transport (Monge and Kantorovich formulations, entropic regularization, Sinkhorn divergences), neural optimal transport estimators, self-supervised image representations (DINOv2), and Cell Painting assays.

Scope in one sentence: The authors formulate molecule-induced phenotype prediction as an inductive conditional transport problem in image-representation space and introduce a molecule-conditioned Neural Optimal Transport model with Monge-Gap regularization that maps negative-control phenotypes toward predicted perturbed phenotypes.

What This Paper Is About

Cell Painting microscopy can profile how cells respond to chemical perturbations, but the chemical space of drug-like molecules is far larger than any imaging campaign can cover, so measuring a phenotype for every candidate molecule is experimentally and economically infeasible. The authors ask whether a model can take a DMSO negative-control phenotype plus the chemical structure of a molecule and predict the perturbed phenotype that molecule would induce. Because control and treated wells are unpaired, they treat the problem as a distributional one and train a molecule-conditioned transport map that generalizes to molecules never seen during training.

Key Contributions

  1. Formulation of phenotype prediction as inductive conditional transport. The authors cast the task as learning a parametric map in phenotypic representation space that transports negative-control (DMSO) phenotypes toward perturbed phenotypes, conditioned on molecular structure — rather than as pointwise regression or static distribution alignment.

  2. A molecule-conditioned Neural OT model with Monge-Gap regularization. Building on the Monge-Gap estimator and its conditional extension, the model injects the molecular condition through multi-head attention, transports autoencoded DINOv2 representations, and incorporates an unbalanced-OT resampling heuristic to down-weight noisy or over-represented regions across experimental batches.

  3. A standardized evaluation protocol using shared controls. Leveraging the eight positive-control molecules repeated across plates in JUMP-CP, the authors define a plate-level (in-distribution) split and a molecule-level (out-of-distribution) split with 5-fold cross-validation and three random seeds, evaluated by cosine-similarity retrieval.

  4. Ablation study isolating what drives generalization. The study identifies the molecular structure encoder as the dominant factor in out-of-distribution performance, shows that compressing the phenotype embedding helps while compressing the conditioning embedding hurts, and characterizes the Monge-Gap weight as governing a trade-off between replicate-level and plate-level retrieval.

Main Findings

  • Static optimal transport does not work here. Classical optimal transport baselines, including the Fused Gromov–Wasserstein instance with feature-cost biasing at known structure–phenotype pairs, yield couplings that are transductive and do not produce useful predictions on large-scale phenotypic image datasets.

  • Neural OT recovers molecule-specific effects across held-out plates. On the in-distribution split, leaving DMSO profiles unchanged (the identity mapping) results in low Retrieval@1. Neural OT improves retrieval for all eight positive-control molecules and substantially increases mean Retrieval@1, reaching or exceeding the held-out-replicate experimental reference while remaining slightly below the plate-mean replicate reference. Molecules with weaker effects remain harder because their target distributions overlap more strongly with DMSO controls and other perturbations.

  • Model-generated phenotypes are batch-robust. In cross-plate retrieval, generated phenotypes achieved a mean average precision of 0.65, compared with 0.36 for raw measured phenotypes and 0.47 after DMSO-based per-plate normalization.

  • Neural OT generalizes to unseen molecules. On the molecule-level out-of-distribution split, NOT reaches R@10 plate = 0.095 ± 0.007, above the identity map (0.059 ± 0.003) and the plate-mean predictor (0.004); the same ordering holds at replicate level (0.035 vs. 0.018 vs. 0.004). The plate-mean predictor is a constant map sitting exactly at the chance level k/M = 10/2248 ≈ 0.004.

  • The molecular encoder is the dominant factor. Under an otherwise identical NOT, retrieval improves markedly from unimol2 and molformer to the fingerprint-based morgan, and is best with the combined morganc+rdkc descriptor, with R@10 plate rising from 0.029 to 0.095. The authors did not evaluate a large pretrained graph encoder in this study.

  • Compressing the condition hurts; compressing the phenotype helps. Reducing the conditioning embedding drops plate R@10 from 0.095 to 0.036–0.037 with an autoencoder or PCA. Reducing the phenotype embedding to 50D by autoencoder reaches 0.101, ahead of PCA (0.075) and no reduction (0.066).

  • Monge-Gap strength trades replicate against plate retrieval. Increasing λ improves replicate retrieval (best at λ = 10, R@10 rep = 0.061) but degrades plate retrieval, which is highest at low regularization (λ = 0, R@10 plate = 0.101; λ = 0.1 gives 0.095). The authors retain λ = 0.1 for a mild geometry-preserving penalty.

  • Attention conditioning beats concatenation. Injecting molecular structure through multi-head attention outperforms plain concatenation (0.095 vs. 0.084 at plate level; 0.035 vs. 0.029 at replicate level).

  • Several design choices matter little. Cosine and squared-Euclidean ground costs perform comparably with overlapping error bars; retrieval is stable across a broad range of the entropic scales; unbalanced resampling gives only small, consistent plate gains around τ ≈ 0.90–0.95.

Methodology in Plain English

The researchers start from Cell Painting images of U2OS cells from the JUMP-CP dataset. Each five-channel microscopy image is converted into an image-level representation using DINOv2, and representations from the same experimental well are averaged into one well-level phenotype — so each well is one replicate of a perturbation. DMSO wells provide the unperturbed "source" phenotypes.

The model learns a map that takes an individual DMSO phenotype plus a representation of a molecule's chemical structure and outputs a predicted perturbed phenotype. The molecule's structure is encoded (as a fingerprint, physicochemical descriptors, or a pretrained molecular embedding) and projected to a common dimension; the control phenotype and the molecular condition are combined through multi-head attention and a feed-forward residual block.

Training is distributional rather than pointwise, because control and treated wells are not paired. The loss combines two terms: a fitting term that pushes the predicted phenotype distribution toward the observed perturbed distribution, and a Monge-Gap regularizer that penalizes the map for being far from an optimal transport plan — using the debiased Sinkhorn divergence and a cosine ground cost matched to the geometry of DINOv2 embeddings. An unbalanced-OT resampling heuristic is optionally applied per mini-batch to reduce the influence of noisy or over-represented samples across plates and batches.

Evaluation is framed as retrieval: predicted phenotypes are ranked against measured candidates by cosine similarity, and retrieval succeeds if the candidate belonging to the correct molecule appears in the top k. The plate-level split holds out plates while keeping the same eight positive-control molecules (measuring generalization across plates and batch effects), and the molecule-level split holds out molecule identities (measuring structure-driven generalization). The out-of-distribution benchmark is restricted to the 12,400 molecules in the top 10% of phenotypic activity, ranked by distance to matched DMSO controls. Training splits in the plate-level setting contain on average 260,262 DMSO points and 142,196 non-DMSO positive-control points, with validation splits of 65,065 DMSO and 35,549 non-DMSO points; the molecule-level setting uses on average 267,785 DMSO source points and 265,496 non-DMSO perturbation points in training, with 263,439 DMSO and 66,374 non-DMSO points in validation, and roughly 30 points per molecule on average. Each fold is trained with three random seeds.

Why This Matters

The work addresses a concrete bottleneck in imaging-based drug discovery: experimental campaigns cannot cover the chemical space of drug-like molecules, so a model that predicts phenotypic responses from structure plus a cheap negative control could prioritize which compounds are worth imaging at all. It also shows that model-generated profiles can serve as batch-robust surrogate representations, which matters because batch effects are a persistent obstacle in comparing morphology measurements across laboratories and plates.

Real-world applications:

  • Compound prioritization and virtual screening. Ranking candidate molecules by predicted phenotype before committing microscopy resources.
  • Batch-robust cross-plate comparison. Using generated profiles as standardized surrogates so measurements from different plates or sites can be compared on equal footing.
  • Hit triage in high-content screening. Filtering large chemical libraries toward compounds likely to induce distinct, reproducible morphological responses.
  • Guiding follow-up imaging campaigns. Deciding which perturbations to acquire experimentally, given that exhaustive characterization is infeasible.

Industry relevance: The authors' affiliations span academic labs and Iktos, a company working in the pharmaceutical space. The framing around chemical space, activity cliffs, and molecule prioritization speaks directly to pharma and biotech workflows that rely on Cell Painting–style phenotypic screening.

Future Directions

  • Stronger molecular representations. The authors identify the molecular encoder as the main limitation on out-of-distribution generalization and suggest stronger molecular graph representations, or training/fine-tuning the molecular encoder during Neural-OT training, to raise the conditioning ceiling. They note that the potentially strongest pretrained graph encoder was unevaluated because it is closed-source.

  • Extension to genetic perturbations. The same conditional-transport formulation could replace the molecular encoder with a gene or protein encoder such as Geneformer or ESM-2, potentially enabling prediction of knockout or over-expression phenotypes.

  • Decoding back to image space. Coupling transported embeddings with a decoder would turn distributional predictions into inspectable morphologies, connecting this approach to generative phenotype models. As it stands, the model predicts phenotype embeddings rather than images and does not by itself yield inspectable morphologies.

  • Harder generalization splits and dispersion analysis. The current split holds out molecule identities rather than molecular scaffolds, so it does not guarantee generalization to structurally distant chemical series. The authors also propose verifying their conjecture about intramolecular prediction dispersion by examining how similar predictions for the same molecule become under different Monge-Gap strengths.

Target Audience

Researchers at the intersection of machine learning and cellular imaging — particularly those working on perturbation response prediction, neural optimal transport, or representation learning for high-content microscopy. It is also relevant to computational drug discovery scientists and screening groups interested in prioritizing compounds, and to machine learning methodologists interested in inductive versus transductive transport formulations and conditional Monge-Gap objectives. The paper is not beginner-friendly: it presumes comfort with optimal transport theory and self-supervised image representations.

Authors’ abstract

High-content microscopy enables systematic profiling of cellular responses to chemical perturbations, but the scale of the chemical space makes exhaustive phenotypic characterization experimentally infeasible. This motivates computational models that can predict image-derived phenotypes without acquiring the corresponding treated cells. We formulate molecule-induced phenotype prediction as an inductive conditional transport problem in image representation space. Given a negative-control phenotype and the structure of a molecule, we aim to predict the phenotype induced by the corresponding molecule. We first evaluate classical optimal transport baselines and show that static couplings do not yield useful predictions on large-scale phenotypic image datasets. We then introduce a molecule-conditioned Neural Optimal Transport (NOT) model with a Monge-Gap regularization training objective that learns to transport negative-control unperturbed phenotypes toward perturbed phenotypes using molecular structure as conditioning information. NOT recovers molecule-specific phenotypic effects while reducing microscopy-associated technical variation, thereby facilitating comparisons across experimental batches. On unseen active molecules, the model outperforms baseline approaches, demonstrating that chemically conditioned transport can generalize beyond the molecules observed during training. We identified the molecular encoder as the main limitation to this generalization, while transport in a compressed representation space improves performance and scalability. These results establish NOT as a promising framework for predicting cellular phenotypes from molecular structure and negative-control phenotypes, while highlighting the development of more informative molecular representations as a key direction for improving out-of-distribution performance.

Read the original paper