Research
VCLMU: Mechanism-Centric Virtual Cell World Modeling for Perturbation Response
Overview Research area: Artificial intelligence for biology and drug discovery — specifically "virtual cell" models that predict how a cell's gene expression changes when a gene is perturbed. The work

- arXiv
- 2610.04475
- Published
- 2026-10-03
- Authors
- Yuwei Miao, Azim Dehghani Amirabad, Scott Oloff, Junzhou Huang, Tianyu Cui, Rui Liao
AI summary
Overview
Research area: Artificial intelligence for biology and drug discovery — specifically "virtual cell" models that predict how a cell's gene expression changes when a gene is perturbed. The work sits at the intersection of generative modeling (world models, latent-variable models) and functional genomics (single-cell CRISPR perturbation screens).
Technical level: Advanced. The method builds on variational latent-variable modeling, object-centric/slot-based representation learning, sparse expert routing, set encoders, and cross-attention decoders. Readers without background in deep generative models or single-cell transcriptomics will find the machinery dense, although the core idea is describable in plain terms.
Scope: This paper proposes VCLMU, a mechanism-centric world model that predicts transcriptional responses to genetic perturbations by treating a perturbation as an action that transforms a set of latent, biologically grounded "mechanism" states, and evaluates it on six held-out perturbation prediction benchmarks.
What This Paper Is About
Most existing perturbation predictors are observation-centric: they take an unperturbed gene-expression profile plus a perturbation and directly output the perturbed profile, without ever explicitly representing the intermediate cellular state change the perturbation causes. The authors argue that scaling this direct mapping has yielded limited returns — they cite a recent systematic evaluation finding that deep perturbation predictors still do not beat simple linear or condition-agnostic baselines.
The goal of VCLMU is to make the latent transition explicit: encode the cell into a set of reusable "mechanism" units, apply the perturbation as an action on those units to produce mechanism-specific state changes, and decode the changed states into the predicted transcriptional response. The paper's bet is that modeling the transition — not just the outcome — improves recovery of perturbation-specific response genes while remaining interpretable.
Key Contributions
-
A world-model formulation of perturbation response. The paper reframes perturbation prediction as action-conditioned latent-state transition in addition to direct observation regression, rather than as a single map from (control profile, perturbation) to response.
-
Latent Mechanism Units (LMUs). A factorized cellular state in which each unit separates a reusable identity (a learnable query blended with a set-encoder summary of a periodically refreshed candidate gene set grounded in multimodal biological evidence) from an observation-specific state (current activity of that mechanism). Under a single shared perturbation action, different active LMUs can undergo different stochastic dynamics.
-
Direct supervision of the latent transition. The observed perturbed profile is encoded through the same control-selected LMUs, and predicted mechanism changes are aligned to observed changes with SmoothL1 using stop-gradient targets. Additional terms — reconstruction of the observed profile from the observed mechanism states, and an inverse model that recovers the perturbed gene from the encoded state change — prevent the target representation from collapsing. Routing and forward dynamics depend only on the control profile and perturbation, so the observed outcome cannot leak into the forward path.
-
A two-stage training curriculum plus interpretability analysis. Pretraining first on approximately 200K pseudo-bulk perturbation profiles, then continued on gene-aligned single-cell perturbation data, with enrichment analysis showing learned LMUs correspond to coherent biological response programs.
Main Findings
-
Best Jaccard@20 on all six datasets. VCLMU-FT (the Stage-2 pretrained model fine-tuned on each evaluation dataset's downstream training split) attains the best Jaccard@20 on all six datasets in Figure 2a.
-
Margin depends on the dataset. The improvement is small on Adamson and Norman, where the condition-agnostic "Perturbed Mean" baseline is already competitive — consistent with reports that simple baselines remain hard to beat. Gains are largest on Replogle K562, Replogle RPE1, and Tian CRISPRi, where baselines recover few true response genes.
-
Pseudo-bulk pretraining transfers. Using the Stage-2 pretrained checkpoint for downstream fine-tuning improves Jaccard@20 on five of six datasets, with the largest gains on Adamson, Norman, and Tian CRISPRa, while preserving performance on Replogle RPE1. The authors interpret this as evidence that pseudo-bulk pretraining captures transferable perturbation structure that single-cell pretraining further sharpens.
-
No trade-off between DE-focused and global accuracy. VCLMU occupies the upper-right region of the ΔPCC / ΔPCC@20 plane, meaning accuracy on top responsive genes does not come at the cost of the global response correlation.
-
Learned LMUs align with real biological programs. Among 1,101 held-out perturbations with at least ten observations, after excluding ubiquitous response genes and subtracting the shared response, 86.6% are significantly associated with at least one LMU at BH-FDR < 0.05. In total, 56 distinct LMUs appear as the best match, with the most frequent accounting for only 20.9% of perturbations. The median best-LMU overlap is seven genes, corresponding to a 19.9-fold enrichment over the random-overlap expectation.
-
LMU-associated programs are coherent and varied. Representative associations span interferon response, mitotic regulation, T-cell signaling, T-cell activation, oxidative phosphorylation, and RNA processing.
-
LMUs are largely non-redundant. Across 50,000 randomly sampled LMU pairs, the median candidate-gene Jaccard is 0.0079, with 47.9% sharing no genes and 80.1% having Jaccard < 0.05.
-
Not reported in the provided content: the specific numeric Jaccard@20, ΔPCC, and ΔPCC@20 values, and the concrete settings of the bank size K, the active-unit count K_a, and the per-LMU candidate-gene count C.
Methodology in Plain English
The model treats a cell like a scene in a world model, with three steps: encode, transition, decode.
Encoding. Instead of one vector per cell, a cell is described by a sparse set of Latent Mechanism Units chosen from a larger bank. Each unit has two parts. Its identity says which biological mechanism it stands for, and is anchored to real biological evidence about a small set of candidate genes the unit "owns" — evidence drawn from functional annotation, pathways, interaction networks, protein sequence, and tissue-expression resources. Its state says how active that mechanism currently is in this particular cell, computed from the expression of those same candidate genes. Which units are active is decided only from the control (unperturbed) profile, so the active set is a property of the starting cell, not the outcome.
Transition. The perturbation is encoded once, from the target gene plus its gene-regulatory-network neighborhood, into an action vector. That same action is applied separately to every active unit. Because each unit starts from a different state, one perturbation can strongly move some mechanisms and leave others untouched. Each unit's update is residual (it predicts a change, not a whole new state) and stochastic — it draws a Gaussian latent variable, regularized toward a standard normal prior by a KL penalty, so the model can represent variability not determined by the control state and the action.
Decoding. The transitioned units are first mixed by a lightweight interaction layer, then a cross-attention decoder lets each gene attend to the predicted mechanism states using its own biological evidence as the query. This produces a per-gene mean and variance defining a Gaussian distribution over the response; the predicted perturbed profile is the control profile plus the predicted mean.
Training signal. A plain reconstruction loss alone would leave the latent transition underdetermined, since many latent paths can explain the same output. So the objective adds mechanism-level supervision: the observed perturbed profile is pushed through the same active units to get observed state changes, and predicted changes are matched to observed changes with a smooth L1 loss using stop-gradient (the encoder supplies the target but is not trained to make that target easy to hit). Two guard terms keep this from degenerating — the observed mechanism states must reconstruct the observed profile, and an inverse model must recover which gene was perturbed from the encoded state change, its difference, and their elementwise interaction. Weak regularizers keep units distinct, anchored, and neither collapsed nor uniformly used.
Two-stage curriculum. Stage 1 pretrains on roughly 200K pseudo-bulk perturbation profiles, built by randomly aggregating 200 cells from the same perturbation condition within each cellular context, drawn from public resources (X-Atlas-Orion, Marson, and the Virtual Cell Challenge 2025 collection) covering more than 10 million single-cell perturbation profiles. Only gene perturbations (CRISPR knockout, interference, and activation) are used; chemical perturbations are excluded. Stage 2 restores that checkpoint, recomputes evidence quantities, freezes the evidence projections, and continues pretraining on gene-aligned single-cell data at a lower learning rate, with the dynamics-alignment and target-state-grounding terms disabled. At inference, the perturbed profile is absent entirely — the model selects units from the control profile, encodes the perturbation, uses the latent mean, and outputs a predicted profile.
Why This Matters
Impact on research. The paper argues that the bottleneck in perturbation prediction may not be model capacity but the failure to model intervention-induced state changes explicitly. If that framing holds, it redirects effort from scaling direct regression toward structured latent dynamics. The interpretability results add a second payoff: the latent units are not opaque — they line up with coherent, largely disjoint transcriptional programs, which makes the model's internal representation inspectable rather than purely predictive. The finding that pseudo-bulk pretraining transfers to single-cell fine-tuning also suggests a practical use for the large pseudo-bulk corpora that already exist.
Real-world applications (implied by the paper's framing):
- Virtual pre-screening of genetic interventions. Before running an expensive CRISPR screen, computationally rank which perturbations are likely to produce a strong, specific transcriptional response.
- Target hypothesis generation for drug discovery. Identify which mechanisms a proposed target perturbation would move, using the LMU structure as a readout of affected biological programs.
- Mechanism interpretation of hits. Use LMU–perturbation associations (interferon response, mitotic regulation, T-cell signaling, oxidative phosphorylation, RNA processing) to assign screen hits to coherent biological programs.
- Reducing wet-lab iterations. The paper frames virtual cell modeling as a step toward "computational evaluation of biological interventions," which in a drug-discovery context supports prioritizing experiments before committing bench resources.
Industry relevance. The work comes from Johnson & Johnson Innovative Medicine with University of Texas at Arlington collaborators and was presented in a workshop titled "AI for Drug Discovery: Bridging the Translation Gap." That framing — plus the emphasis on not sacrificing global response accuracy while improving response-gene recovery — speaks directly to industrial teams that need perturbation predictions to be both accurate and mechanistically legible before they can inform target selection or experimental design.
Future Directions
-
Extending LMU validation to multimodal settings. The conclusion explicitly states this as future work: the current validation is transcriptomic, and the authors plan to test whether LMUs remain biologically coherent when other modalities are involved.
-
Combinatorial perturbations. The paper notes that the current study focuses primarily on single-gene transcriptomic perturbations and that extending to combinatorial perturbation settings is planned. Whether the shared-action, per-unit transition mechanism scales to multiple simultaneous actions is an open question.
-
Closing the gap on datasets where simple baselines remain competitive. VCLMU's margin is small on Adamson and Norman, where the condition-agnostic Perturbed Mean baseline is already strong. Improving on those regimes — rather than only on datasets where baselines recover few response genes — remains unresolved.
-
Better characterizing what the latent transition variable captures. The stochastic per-unit latent is regularized toward a standard normal prior, and the authors note the forward path deliberately excludes outcome information. What biological variability that latent actually absorbs — versus what is determined by control state and action — is not established by the analyses presented.
Target Audience
- Machine learning researchers working on generative and world models who are interested in object-centric or slot-based latent representations applied to a non-standard domain, and in supervising latent transitions rather than only outputs.
- Computational biologists and bioinformaticians working on single-cell perturbation screens, virtual cell models, and benchmark-driven evaluation of perturbation predictors.
- Drug-discovery scientists in industry — particularly those in target identification, functional genomics, and computational biology groups — who need perturbation models whose predictions are interpretable enough to inform experimental decisions.
- Benchmark and evaluation researchers interested in perturbation-disjoint splits, the Jaccard@20 / ΔPCC / ΔPCC@20 evaluation setup, and the recurring problem that simple baselines are hard to beat.
Readers need prior familiarity with variational latent-variable models and single-cell expression data to follow the method section in detail; the abstract, introduction, results narrative, and interpretability findings are accessible to a broader scientific audience.
Authors’ abstract
Predicting cellular responses to genetic perturbations is a central capability for virtual cells and a key step toward computational modeling of biological interventions. Most existing models directly map an unperturbed molecular profile and perturba- tion to the resulting observation without explicitly representing the latent cellular transition induced by the intervention. We introduce a mechanism-centric virtual cell world model that represents cellular state as a set of Latent Mechanism Units (LMUs) and treats genetic perturbations as actions on these latent states. Each LMU combines a reusable identity grounded in multimodal biological evidence with an observation-specific state, allowing a perturbation to induce mechanism- specific stochastic transitions before decoding the resulting transcriptional response. We train VCLMU through two-stage pretraining, first on around 200K pseudo-bulk perturbation profiles and then on gene-aligned single-cell perturbation data. Across six perturbation-disjoint benchmarks, VCLMU consistently improves perturbation- specific response recovery over strong baselines while maintaining competitive global response accuracy. We further analyze learned LMUs through enrichment between perturbation responses and LMU gene sets and show that they capture structured biological response programs. These results support mechanism-level latent state transition as a useful formulation for virtual cell models that aim to predict and interpret cellular responses to biological interventions.