Research
Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
Overview Research area: Machine learning / AI for science — specifically the role of foundation models (FMs) in scientific discovery. Technical level: Beginner-Friendly. This is a position paper (arXi
- arXiv
- 2510.15280
- Published
- 2025-10-17
- Authors
- Fan Liu, Jindong Han, Tengfei Lyu, Weijia Zhang, Zhe-Rui Yang, Lu Dai, Cancheng Liu, Hao Liu
AI summary
Overview
Research area: Machine learning / AI for science — specifically the role of foundation models (FMs) in scientific discovery.
Technical level: Beginner-Friendly. This is a position paper (arXiv:2510.15280v1 [cs.LG], 17 Oct 2025) that advances a conceptual framework and surveys existing FM applications rather than presenting new experiments or benchmarks.
Scope in one sentence: The paper argues that foundation models such as GPT-4 and AlphaFold are not merely accelerating existing scientific methods but are driving a transition toward a new, fifth scientific paradigm, and it organizes that transition into a three-stage framework spanning meta-scientific integration, hybrid human-AI co-creation, and autonomous scientific discovery.
What This Paper Is About
Scientific practice has traditionally advanced through four paradigms: experiment-driven, theory-driven, computation-driven, and data-driven science. As problems such as consciousness, protein folding pathways, social polarization, drug discovery, and materials design resist reductionist modeling or involve combinatorial explosion, the limits of these paradigms become more visible. The paper asks whether foundation models are simply enhancing existing scientific methodologies or redefining how science is conducted, and argues the latter — proposing a three-stage framework to describe the shift.
Key Contributions
-
A new conceptual framework. A three-stage model of FM-driven scientific evolution — Meta-Scientific Integration, Hybrid Human-AI Co-Creation, and Autonomous Scientific Discovery — supported by a five-dimension comparison table (paradigm definition, FM role, task scope, autonomy, and impact on science).
-
A systematic review and taxonomy. An analysis of FM-enabled scientific discovery organized by how FMs integrate into experimental, theoretical, computational, and data-driven workflows, plus a section on cross-paradigm integration.
-
A research agenda. Identification of key risks that must be addressed to realize the scientific potential of FMs, and directions for future research on aligning epistemic goals with emerging AI capabilities.
-
A public resource. The authors release an accompanying project repository at https://github.com/usail-hkust/Awesome-Foundation-Models-for-Scientific-Discovery.
Main Findings
-
FMs as infrastructure, not yet as epistemically transformative agents (Stage 1). In meta-scientific integration, FMs automate data preprocessing, literature retrieval, and methodology matching, and link previously isolated components such as sensor data with simulation models. They remain bound by human-defined objectives, exhibit low autonomy, and their role is instrumental rather than epistemic — the locus of reasoning stays human.
-
FMs are becoming active collaborators (Stage 2). In hybrid human-AI co-creation, FMs contribute to research question generation, hypothesis structuring, experiment planning, execution support, interpretation, and scientific discourse. They exhibit moderate autonomy and can influence the trajectory of discovery, but human prompts still frame problems and provide ethical guidance. The paradigm reshapes the division of cognitive labor without changing who conducts science.
-
Autonomous scientific discovery is framed as a fifth paradigm (Stage 3). The paper envisions FMs that pose questions, generate hypotheses, select methods, run experiments or simulations, interpret results, and update internal models with minimal human oversight. Systems like AI scientists are described as having conducted the whole research pipeline, and FunSearch is highlighted as autonomously proposing and validating new mathematical conjectures.
-
Concrete integration exists across all four classical paradigms. Experiment-driven: FMs as priors or feature extractors in Bayesian optimization and active learning pipelines, and as planners generating Python control scripts for instruments (with examples including LLM-RDF, CLAIRify, VISION, and AP-VLM). Theory-driven: knowledge-graph-guided hypothesis generation (e.g., KG-CoI, the HypoGen dataset) and neuro-symbolic validation (e.g., Logic-LM, Popper, LeanCopilot, DeepSeekProver). Computation-driven: symbolic discovery (LLM-SR, FunSearch), latent operators (PROSE-PDE, DiffusionPDE), and neural operators and weather models (GraphCast, FactFormer, PDE-Refiner). Data-driven: DNABERT, MoLFormer, ChemVLM, ClimaX, Galactica, Pangu-Weather, DiffusionSat, AlphaFold 2, ESMFold, RFdiffusion, and MatterGen.
-
Cross-paradigm integration is emerging. The paper points to PROSE-FD, which co-trains symbolic equation templates and spatial field data in a multimodal Transformer, Latent Neural Operators (LNOs) for geometry-agnostic forward and inverse problems, and Coscientist, which translates high-level research goals into machine-executable protocols and controls robotic synthesis.
-
Four risk dimensions intensify with autonomy. (1) Bias and epistemic fairness — FMs inherit biases from training data that overrepresent dominant paradigms, Western institutions, and widely cited authors; the paper's example is global health modeling prioritizing diseases well-studied in Western contexts while overlooking issues such as schistosomiasis or child stunting in sub-Saharan Africa. (2) Hallucination and scientific misinformation — FMs remain pattern recognizers, not truth-preserving reasoners. (3) Reproducibility and scientific transparency — opaque reasoning threatens replication. (4) Authorship, accountability, and scientific ethics — questions of credit, ghost authorship, and misuse arise as FMs move toward autonomy.
-
Current FM deployments remain limited. Most are confined to static prompts, predefined tasks, and fixed schema, and typically lack persistent memory, adaptive feedback, and physical embodiment, making contributions largely reactive and limited to isolated stages.
-
No quantitative benchmarks are reported. As a position paper, it contains no experimental results, dataset sizes, or performance metrics beyond qualitative references to other systems; CLIP is described as trained on 400 million image–text pairs.
Methodology in Plain English
The authors take a conceptual and review-based approach rather than an experimental one. They begin by recounting the history of scientific discovery as four successive paradigms — experiment-driven (16th–17th century, Galileo and Boyle), theory-driven (18th–19th century, Newton, Maxwell, Einstein), computation-driven (mid-20th century), and data-driven (21st century). They then position foundation models against this history, characterizing FMs as large-scale neural networks trained on massive, diverse datasets through unsupervised or self-supervised pretraining followed by fine-tuning or prompting. From this, they propose the three-stage framework and build a comparison table across five dimensions. The body of the paper then surveys existing FM systems, sorting them by the classical paradigm they touch, and adds a section on systems that bridge multiple paradigms. Finally, they conduct a risk analysis and propose future research directions, concluding that these developments point toward a "fifth scientific paradigm."
Why This Matters
Impact on research. The paper reframes debates about AI in science from a tooling question ("does this speed things up?") to an epistemic one ("who or what produces scientific knowledge?"). It gives researchers and policymakers a shared vocabulary — meta-scientific integration, hybrid human-AI co-creation, autonomous scientific discovery — for discussing where a given system sits on the autonomy spectrum and what governance it needs.
Real-world applications discussed in the paper:
- Drug discovery and materials design, where FMs navigate combinatorial explosions of candidate spaces that make exhaustive search infeasible.
- Global health modeling, where bias in FM training data can skew which diseases receive research attention.
- Automated laboratory work, with FMs generating instrument control scripts and coordinating robotic synthesis (LLM-RDF, CLAIRify, Coscientist).
- Weather, climate, and Earth-system science, through models such as GraphCast, Pangu-Weather, and ClimaX that learn latent spatio-temporal dynamics at reduced computational cost.
- Protein structure prediction and design, via AlphaFold 2, ESMFold, and RFdiffusion.
Industry relevance. The framework is directly relevant to companies building laboratory automation, AI-driven drug and materials discovery platforms, and scientific computing tools, because it maps specific capability levels onto specific accountability and validation requirements. The risk section — bias, hallucination, reproducibility, and authorship — speaks to regulatory and governance concerns that any organization deploying FMs in a research or R&D pipeline will face.
Future Directions
-
Embodied scientific agents. Grounding FMs in laboratory robotics, automated instruments, and digital twin environments so they can plan experiments, interact with physical systems, and iteratively refine procedures. Open challenges include integrating high-level task planning with low-level control, robustness under real-world uncertainty, and safety and interpretability in dynamic lab environments.
-
Closed-loop scientific autonomy. Moving from open-loop workflows, where humans decide next steps, to systems where FMs continuously formulate hypotheses, run experiments, analyze results, and update internal models. The paper cites reinforcement learning-based planning, planning-as-inference, and neuro-symbolic agents as early progress, and flags the challenge of keeping the loop robust to noisy observations, adaptive to shifting objectives, and aligned with scientific validity rather than mere reward maximization.
-
Continual learning and generalization. Enabling FMs to accumulate and refine knowledge over time, addressing catastrophic forgetting and domain drift through parameter-efficient online adaptation, memory-augmented architectures, and modular lifelong learning frameworks. The stated goal is incremental construction of domain-bridging representations, analogical reasoning across scientific contexts, and coherent long-horizon research trajectories.
-
Governance and validation. The paper calls for diverse training data, fairness-aware evaluation, verification mechanisms such as symbolic logic checks and simulation-based validation, transparent logging of reasoning steps, version-controlled model checkpoints, and governance frameworks that distinguish mechanical from creative contributions and mandate disclosure.
Target Audience
This paper is best suited to researchers, graduate students, and scientific administrators who want a conceptual map of how foundation models are entering scientific practice. It is particularly useful for scientists in computationally intensive fields (chemistry, materials science, biology, climate science, mathematics) seeking orientation on FM capabilities, for AI researchers interested in the epistemology of autonomous discovery, and for policymakers and research-integrity officers thinking about bias, reproducibility, and authorship in AI-assisted science. Readers looking for experimental results, benchmarks, or implementation details will not find them here — the value lies in the framework, the taxonomy, and the risk agenda.
Authors’ abstract
Foundation models (FMs), such as GPT-4 and AlphaFold, are reshaping the landscape of scientific research. Beyond accelerating tasks such as hypothesis generation, experimental design, and result interpretation, they prompt a more fundamental question: Are FMs merely enhancing existing scientific methodologies, or are they redefining the way science is conducted? In this paper, we argue that FMs are catalyzing a transition toward a new scientific paradigm. We introduce a three-stage framework to describe this evolution: (1) Meta-Scientific Integration, where FMs enhance workflows within traditional paradigms; (2) Hybrid Human-AI Co-Creation, where FMs become active collaborators in problem formulation, reasoning, and discovery; and (3) Autonomous Scientific Discovery, where FMs operate as independent agents capable of generating new scientific knowledge with minimal human intervention. Through this lens, we review current applications and emerging capabilities of FMs across existing scientific paradigms. We further identify risks and future directions for FM-enabled scientific discovery. This position paper aims to support the scientific community in understanding the transformative role of FMs and to foster reflection on the future of scientific discovery. Our project is available at https://github.com/usail-hkust/Awesome-Foundation-Models-for-Scientific-Discovery.