Research
Universally Converging Representations of Matter Across Scientific Foundation Models
Universally Converging Representations of Matter Across Scientific Foundation Models Overview Research area: Machine learning for science — representation learning and foundation models for molecules,

- arXiv
- 2512.03750
- Published
- 2025-12-03
- Authors
- Sathya Edamadaka, Soojung Yang, Ju Li, Rafael Gómez-Bombarelli
AI summary
Universally Converging Representations of Matter Across Scientific Foundation ModelsOverview
- Research area: Machine learning for science — representation learning and foundation models for molecules, materials, and proteins.
- Technical level: Intermediate (assumes familiarity with embeddings, latent spaces, and ML interatomic potentials, but the core ideas are explained conceptually).
- Scope: A representation-level comparison of 59 scientific machine learning models spanning string, graph, 3D atomistic, protein, and natural-language modalities, testing whether they learn a shared latent representation of matter.
What This Paper Is About
Machine learning models for molecules, materials, and proteins use very different inputs and architectures, but all are ultimately trying to learn the same underlying physics. This paper asks whether those models actually learn similar internal representations — and whether that similarity grows as models get better. The authors use representational alignment both as a scientific question and as a benchmark for deciding whether a model is genuinely "foundational" rather than merely accurate on its training domain.
Key Contributions
- A large-scale cross-modal alignment study. The authors measure representational similarity across 59 scientific models spanning string encodings (SMILES/SELFIES), 3D atomic coordinates, protein sequences, protein structures, and natural language, using five datasets (QM9, OMol25, OMat24, sAlex, and RCSB).
- Evidence that representations converge with performance. They show that machine learning interatomic potentials grow more aligned with top-performing models as their energy prediction error decreases, and that intrinsic dimensionality is consistent across architectures within a dataset.
- A two-regime failure taxonomy. They identify that on in-distribution structures, weak models diverge into many local sub-optima, while on out-of-distribution structures nearly all models collapse onto low-information, architecture-specific manifolds.
- Representational alignment as a generality benchmark. They propose alignment with other high-performing models as a quantitative criterion for foundation-level generality, and use it to draw practical lessons about data diversity, model distillation, and architectural choices such as the Orb V3
equigradregularization scheme.
Main Findings
- Models align strongly even across modalities. Representations from string-based molecule models and 3D-atomistic MLIPs are significantly aligned. Although these cross-modality alignment values are lower than within-modality values, they exceed the highest representation alignment values previously found between language and vision foundation models.
- Protein models align most strongly. Protein sequence models and protein structure models align nearly twice as strongly as the best cases for small molecules, consistent with evidence that large protein sequence models implicitly learn folding constraints and structural regularities.
- Large language models participate in the alignment. General-purpose LLMs given SMILES strings align strongly with string-based materials models and show alignment scores with MLIPs comparable to those of other SMILES-based models.
- Representations converge as performance improves. As models decrease in energy regression MAE on materials from OMat24, they become more aligned with the best-performing model. For in-distribution materials, UMA Medium emerges as the convergence target; for QM9 small molecules, models converge toward an Orb V3 Conservative model.
- Local and global metrics agree. CKNNA (a local, neighborhood-based metric) is highly correlated with distance correlation (a purely global metric) for small-molecule representations, and CKNNA values remain stable as the degree of locality is increased.
- Intrinsic dimensionality is remarkably consistent. QM9 representations have an intrinsic dimensionality of roughly 5, while OMat24 (roughly 10), sAlex (roughly 8), and OMol25 (roughly 10) require more dimensions — differences the authors attribute to the diversity of chemical environments sampled in each dataset. Equivariant models have higher intrinsic dimensionality than invariant embeddings, and "vanilla" models have consistently higher values than both.
- Training data shapes representations more than architecture — in-distribution. An evolutionary tree built from CKNNA distances clusters models by both architecture and training dataset, but MACE trained on the small-molecule OFF dataset diverges from other MLIPs including the architecturally similar MACE-MP. eSEN models trained on OMat24 align more closely with EqV2 models trained on OMat24 than with eSEN models trained on MPTraj.
- Two distinct failure regimes. On in-distribution data, weakly performing models are weakly aligned with each other and occupy divergent, non-generalizable sub-optima; models cluster by training dataset. On out-of-distribution data (OMol25 structures), almost all models are poorly performing yet learn very similar information, collapsing toward architecture-specific manifolds — this time clustering by architecture rather than dataset. Latent spaces are therefore data-dominated in-distribution and architecture-dominated out-of-distribution.
- A concrete example of a local optimum. MACE-OFF achieves strong performance on QM9 but shows weak alignment to the Orb V3 family, UMA, and other high-performing models. The authors note Orb V3 substantially outperforms MACE-OFF on the GMTKN55 benchmark, a broader and more chemically diverse molecular dataset than QM9.
- Regularization can substitute for architectural equivariance. Orb V3 conservative models align strongly with fully equivariant architectures like MACE and EqV2, whereas the direct Orb V3 variant without
equigradaligns weakly — suggesting an appropriately structured regularization scheme can reproduce the representational benefits of architectural equivariance at lower computational cost. - Current models are not yet foundational. In-distribution, representations cluster by training dataset; out-of-distribution, they collapse by architecture. The authors conclude that today's models remain limited by training data and inductive bias and do not yet encode truly universal structure.
Methodology in Plain English
The authors took 59 existing scientific models — ranging from small graph neural networks to models with hundreds of billions of parameters — and fed them structures from five standard datasets. For each model, they saved the numerical embedding produced at the last hidden layer before the model's output head, averaging across atoms where models produce per-node embeddings. They sampled 50,000 structures from each dataset for the metric comparisons, and used smaller subsets (for example, 1,000 or 15,000 structures) for specific visualizations and convergence plots.
They then compared these embeddings with four different measures: CKNNA, which looks at whether nearby points in one model's space are also nearby in another's; distance correlation, a global measure of how two spaces covary; intrinsic dimensionality, which estimates how many variables are needed to describe an embedding space; and information imbalance, an asymmetric measure of which representation contains more information than another. To compare models trained at different levels of density functional theory (PBE-OMat24 versus PBE-MP), they fit a linear model of energy versus elemental composition for each and compared only the residual deviations from that compositional baseline. Inference ran mostly on a single 32 GB V100 GPU, with four 80 GB A100 GPUs used for LLM inference. The LLMs were tested under three prompt conditions: an extended system prompt with SMILES strings, a minimal prompt, and a minimal prompt with an ASE Atoms object readout.
Why This Matters
This work reframes model evaluation. Instead of judging scientific models only by prediction error on a benchmark, it asks whether a model's internal representation matches those of other strong models — which distinguishes genuine generalization from overfitting to a narrow chemical space.
Impacts on research:
- Provides a model-agnostic, quantitative benchmark for "foundation-level" generality in scientific models, and a way to track whether universal representations emerge as models scale.
- Shows that many architectures converge to similar solutions, which supports model distillation and reuse of learned embeddings rather than training every model from scratch.
- Reveals that today's materials models are data-limited rather than architecture-limited, pointing to dataset diversity (equilibrium and non-equilibrium regimes, broader chemical and structural environments) as the main bottleneck.
- Explains observed scaling behaviors: models from the same architecture family learn similar representations even at vastly different sizes, so small models can mimic the expressivity of larger ones.
Real-world applications:
- Faster generative model training: the authors cite prior work where adding a loss that promoted alignment with a pre-trained model accelerated generative model training, including in the sciences where a generative model sampling equilibrium conformational ensembles trained faster with an alignment regularization term toward a pre-trained MLIP.
- Property prediction and spectroscopy: MLIP representations have been used for NMR chemical shift prediction and for out-of-domain electronic structure and excited state predictions, producing more physically consistent results than task-specific baselines.
- Efficient model selection: researchers can pick smaller or cheaper architectures whose representations align closely with large, high-performing ones, inheriting expressive representational structure at lower cost.
- Deployment of non-equivariant models: the Orb V3 result suggests that well-designed regularization plus scale can approximate symmetry-enforcing models, which matters for cost-sensitive simulation and inference.
Industry relevance: labs and companies building materials discovery, drug design, and protein engineering pipelines often must choose among dozens of pretrained models and cannot afford to train fully equivariant or conservative models themselves. This paper offers a principled selection criterion and a diagnostic for when a model is likely to fail outside its training distribution.
Future Directions
- Tracking convergence with scale. The authors position representational alignment as a tool for monitoring whether universal representations of matter emerge as models grow larger and are trained on more diverse data.
- Closing the data gap for materials. Their analysis implies materials models need substantially more diverse training data spanning equilibrium (sAlex) and non-equilibrium (OMat24) regimes and a wider range of chemical environments, like those used to train Meta's UMA model.
- Building genuinely out-of-distribution-capable models. Because nearly all models collapse onto low-information, architecture-specific manifolds on unseen structures, a central open question is how to design architectures and training objectives whose inductive biases do not dominate in that regime.
- Using alignment for selection and distillation. Further work could operationalize alignment scores to automatically select models whose representations transfer best across modalities, types of matter, and scientific tasks, and to guide distillation into cheaper architectures.
Target Audience
Machine learning researchers working on scientific foundation models, computational chemists and materials scientists choosing or benchmarking ML interatomic potentials, computational biologists working with protein sequence and structure models, and method developers interested in representation similarity, model distillation, and the empirical study of inductive bias versus data scaling.
Authors’ abstract
Machine learning models of vastly different modalities and architectures are being trained to predict the behavior of molecules, materials, and proteins. However, it remains unclear whether they learn similar internal representations of matter. Understanding their latent structure is essential for building scientific foundation models that generalize reliably beyond their training domains. Although representational convergence has been observed in language and vision, its counterpart in the sciences has not been systematically explored. Here, we show that representations learned by nearly sixty scientific models, spanning string-, graph-, 3D atomistic, and protein-based modalities, are highly aligned across a wide range of chemical systems. Models trained on different datasets have highly similar representations of small molecules, and machine learning interatomic potentials converge in representation space as they improve in performance, suggesting that foundation models learn a common underlying representation of physical reality. We then show two distinct regimes of scientific models: on inputs similar to those seen during training, high-performing models align closely and weak models diverge into local sub-optima in representation space; on vastly different structures from those seen during training, nearly all models collapse onto a low-information representation, indicating that today's models remain limited by training data and inductive bias and do not yet encode truly universal structure. Our findings establish representational alignment as a quantitative benchmark for foundation-level generality in scientific models. More broadly, our work can track the emergence of universal representations of matter as models scale, and for selecting and distilling models whose learned representations transfer best across modalities, domains of matter, and scientific tasks.