Research
Modeling Microenvironment Trajectories on Spatial Transcriptomics with NicheFlow
NicheFlow: Modeling Microenvironment Trajectories on Spatial Transcriptomics Overview Research area: Machine learning for spatial transcriptomics — specifically generative flow-matching models for inf

- arXiv
- 2511.00977
- Published
- 2025-11-02
- Authors
- Kristiyan Sakalyan, Alessandro Palma, Filippo Guerranti, Fabian J. Theis, Stephan Günnemann
AI summary
NicheFlow: Modeling Microenvironment Trajectories on Spatial TranscriptomicsOverview
- Research area: Machine learning for spatial transcriptomics — specifically generative flow-matching models for inferring the temporal evolution of tissue microenvironments across time-resolved spatial slides.
- Technical level: Advanced. The paper assumes familiarity with flow matching, optimal transport (including entropic OT and couplings on incomparable spaces), variational inference, point-cloud transformers, and single-cell/spatial transcriptomics terminology.
- Scope: The paper introduces NicheFlow, a flow-based generative model that treats cellular neighborhoods ("niches") as point clouds and learns to generate their future spatial coordinates and gene-expression states, evaluated on three spatiotemporal datasets spanning embryonic, brain, and ageing biology.
What This Paper Is About
Spatial transcriptomics measures gene expression while preserving each cell's location in tissue, but each slide is only a static snapshot, and because tissue is destroyed during acquisition there is no direct cell-to-cell correspondence between slides. Existing trajectory-inference methods model the movement of individual cells, which misses the coordinated evolution of structured neighborhoods. This paper's goal is to instead model how entire local microenvironments — a cell plus its fixed-radius neighborhood — evolve from one time point to the next, generating both where cells will be and what state they will be in.
Key Contributions
- A microenvironment-centered trajectory inference paradigm. Instead of tracking individual cells over time, NicheFlow models niches as point clouds, predicting spatial coordinates and gene expression profiles jointly while preserving local tissue context.
- A factorized Variational Flow Matching (VFM) approach. The posterior is factorized across individual points and across coordinate and feature dimensions, with different distributional families used for each: Laplace for spatial coordinates and Gaussian for gene expression, trained with a correspondingly factorized loss (an L1 term for coordinates and an L2 term for features).
- A spatially-aware sampling strategy. Source and target microenvironments are matched via entropic optimal transport computed on pooled niche representations, with a tunable hyperparameter λ balancing spatial versus cellular-state information. Mini-batches are drawn uniformly from K-Means-defined regions of the slide to ensure coverage of heterogeneous tissue.
- Empirical validation across diverse spatiotemporal data. The model is tested on mouse embryonic development, axolotl brain development, and mouse brain ageing, reportedly outperforming baselines in recovering cell-type organization and spatial structure.
Main Findings
- Best overall quantitative performance for the proposed configuration: NicheFlow trained with the Gaussian-Laplace VFM objective (GLVFM) achieves the highest 1NN-F1 and lowest PSD on all three datasets among the methods compared (values reported as mean ± standard deviation over five evaluation runs). On mouse embryonic development (MED): 1NN-F1 0.664 ± 0.0014, PSD 0.883 ± 0.0094, SPD 0.398 ± 0.0023. On axolotl brain development (ABD): 1NN-F1 0.628 ± 0.0013, PSD 2.079 ± 0.0043, SPD 0.576 ± 0.0055. On mouse brain ageing (MBA): 1NN-F1 0.285 ± 0.0003, PSD 1.554 ± 0.0021, SPD 0.532 ± 0.0009.
- One marginal exception on a coverage metric: On the ageing dataset, NicheFlow GVFM reports a slightly lower SPD (0.531) than NicheFlow GLVFM (0.532), while GLVFM has the higher 1NN-F1 (0.285 vs 0.268) and lower PSD (1.554 vs 1.661).
- Single-cell modeling is insufficient: SPFlow, the single-point flow baseline, performs worst among the flow-based methods. For example, with CFM it reaches 1NN-F1 0.272 ± 0.0011 on MED, 0.190 ± 0.0005 on ABD, and 0.205 ± 0.0003 on MBA. The authors describe SPFlow as failing to capture tissue-level organization, producing blurry and spatially incoherent samples.
- Spatially co-localized conditioning matters: RPCFlow, which uses the same backbone as NicheFlow but conditions on randomly sampled point clouds rather than radius neighborhoods, is reported to generate diffused mappings across the slide and to fail to preserve spatial consistency. Its best results are with GLVFM (MED 1NN-F1 0.586 ± 0.0016, ABD 0.554 ± 0.0007, MBA 0.265 ± 0.0004), still below NicheFlow.
- LUNA as a reference point only: LUNA, a diffusion model for spatial reconstruction from dissociated cells, does not model temporal dynamics and only generates coordinates from noise with biological annotations. It is used as a reference for spatial generation accuracy via 1NN-F1: 0.540 ± 0.004 on MED, 0.331 ± 0.003 on ABD, and 0.222 ± 0.000 on MBA. No PSD or SPD values are reported for it in the table.
- Objective choice matters for the proposed backbone: NicheFlow with GLVFM outperforms its own CFM variant on 1NN-F1 for all three datasets (MED 0.664 vs 0.609; ABD 0.628 vs 0.604; MBA 0.285 vs 0.283).
- Qualitative biological validation: Using the mouse embryonic dataset, the authors examine two scenarios driven by the choice of λ: compositional changes in fixed structures (using the spinal cord from E10.5 to E11.5, with a low λ to prioritize spatial preservation), and spatial and cellular development of immature cells that may displace to different areas, which requires a higher λ to account for gene expression. The paper content is truncated mid-sentence in this qualitative case study.
- The λ setting used in the main benchmark: All experiments in the main comparison use a fixed λ = 0.1, described as enabling spatial location preservation across time; comparisons across multiple λ values are stated to appear in the paper's Table 3.
Methodology in Plain English
The input to the method is a tissue slide represented as an "attributed point cloud" — each cell is a 2D coordinate paired with a feature vector (gene expression, reduced to its top 50 principal components). To move from single cells to neighborhoods, the authors build a microenvironment for every cell: all cells within a fixed spatial radius r. Source and target collections of these neighborhoods are built separately for each time point.
Because no cell-level correspondence exists across slides, the method matches source and target niches using entropic optimal transport on a pooled representation of each niche — the average coordinates and average expression features, blended by the hyperparameter λ (a value of 0.1 was used for the main experiments, favoring spatial proximity). The resulting coupling supplies training pairs.
The generative model itself is a source-conditioned Variational Flow Matching setup: given noise and a source microenvironment, the model predicts a target microenvironment. Following the variational view of flow matching, only the mean of the posterior over the target is needed when using straight-line interpolation paths, so the model learns a time-dependent predictor of mean coordinates and mean features rather than a full distribution. This mean predictor is parameterized by a permutation-invariant point-cloud transformer with an encoder-decoder layout: the encoder self-attends over the source microenvironment, the decoder self-attends over the noisy target and cross-attends to the encoder output, with time encoded via sinusoidal embeddings and separate embeddings for coordinates and features. Training minimizes a factorized loss combining an L1 penalty on coordinate predictions and a squared L2 penalty on feature predictions.
For training, mini-batches of 64 regions are sampled uniformly from K-Means clusters over 2D coordinates for spatial diversity. For evaluation, each tissue is discretized into a fixed 2D grid, and evaluation microenvironments are defined around the nearest cells to each grid point, which the authors state guarantees full spatial coverage and deterministic comparison across methods. Three metrics are used: PSD (point-to-shape distance, squared distance from each generated point to its nearest ground-truth counterpart), SPD (shape-to-point distance, squared distance from each ground-truth point to its nearest generated point), and 1NN-F1 (a weighted F1 score where generated cells are labeled by a classifier trained on ground-truth expression and matched to nearest real cells). A single model is trained with additional conditioning on source and target time labels so it can predict piecewise trajectories across more than two time points.
Why This Matters
- Shifts the unit of analysis: Most trajectory inference operates on isolated cells. This work argues that coordinated neighborhoods — functionally distinct niches shaped by cell-to-cell interactions and extracellular components — are the more biologically natural unit for understanding tissue-scale processes.
- Tackles a real technical obstacle: The destructive nature of spatial acquisition means there is no cell correspondence between slides; the paper's OT-plus-flow formulation is designed specifically around this missing-correspondence problem, and around the fact that source and target niches can have different numbers of cells.
- Scalability framing: By learning dynamics of variably located sub-regions rather than whole slides, and by using mini-batch deep learning instead of exact OT at the single-cell level, the approach is positioned against methods the authors describe as limited in scalability and generalization.
Real-world applications the paper's framing points toward:
- Developmental biology: Mapping how embryonic tissues such as the spinal cord reorganize across stages (the paper studies E9.5, E10.5, and E11.5 in mouse, and five time points in axolotl brain).
- Ageing research: Using the mouse brain ageing dataset (twenty MERFISH time points) to study how tissue architecture and cell-type organization change with age.
- Cancer and immunology: The introduction cites tumor progression and immune infiltration as processes governed by spatial microenvironments, suggesting trajectory modeling of niches could inform these areas.
- Tissue regeneration: Also cited as a microenvironment-driven process, which time-resolved spatial modeling could help characterize.
Industry relevance: The work sits at the intersection of computational biology tooling and generative modeling, relevant to pharmaceutical and biotech research groups that use spatial transcriptomics for target discovery, to developers of single-cell and spatial analysis platforms, and to machine learning researchers interested in flow matching and optimal transport over structured, variable-size data.
Future Directions
- Extending beyond pairwise and short sequences: The formulation is presented for two consecutive time points (s ∈ {0,1}), with the authors noting it can be extended to more consecutive discrete temporal measurements, and they train a single time-label-conditioned model for multiple time points. How well this scales across long sequences, such as the twenty-time-point ageing dataset, remains an open area.
- Choosing the spatial-versus-state trade-off: The λ hyperparameter governs whether the matching prioritizes spatial location or gene expression, and the authors present it as a scenario-dependent choice rather than something automatically determined. How to set or learn λ in a principled way is left open.
- Modeling limitations: The paper explicitly states in its Section 4.3 that the method has modeling limitations outlined in Appendix B; the provided content does not detail them, so the specific constraints are not reported here.
- Transfer beyond biology: The authors state that principled learning of dynamics in structured, high-dimensional, and variably sized spatial domains has parallels in other spatiotemporal modeling domains beyond biology, suggesting generalization of the point-cloud niche formalism to other structured spatiotemporal data.
Target Audience
This paper benefits most readers with a background in machine learning for biological data — graduate students and researchers in computational biology, bioinformatics, and generative modeling — as well as method developers familiar with flow matching, optimal transport, or point-cloud generative models who are interested in structured, variable-size data. Biologists and clinicians working with time-resolved spatial transcriptomics will find the biological case studies and evaluation metrics relevant, though they may need to skim the variational and OT formalism. Readers seeking a gentle introduction to trajectory inference should look for a more introductory source first, given the paper's reliance on entropic optimal transport and variational flow matching theory.
Authors’ abstract
Understanding the evolution of cellular microenvironments in spatiotemporal data is essential for deciphering tissue development and disease progression. While experimental techniques like spatial transcriptomics now enable high-resolution mapping of tissue organization across space and time, current methods that model cellular evolution operate at the single-cell level, overlooking the coordinated development of cellular states in a tissue. We introduce NicheFlow, a flow-based generative model that infers the temporal trajectory of cellular microenvironments across sequential spatial slides. By representing local cell neighborhoods as point clouds, NicheFlow jointly models the evolution of cell states and spatial coordinates using optimal transport and Variational Flow Matching. Our approach successfully recovers both global spatial architecture and local microenvironment composition across diverse spatiotemporal datasets, from embryonic to brain development.