Research
Toward a Locally Deployable Agentic Co-Scientist: Small-Model Planning for Early-Stage Drug Discovery
Overview Research area: AI for drug discovery — specifically LLM-based agentic workflow orchestration for early-stage computational drug discovery. Technical level: Intermediate. Readers should be com

- arXiv
- 2610.04740
- Published
- 2026-10-03
- Authors
- Tian Liang, Jiayu Chang, Alejandro F. Frangi, Mobarak I. Hoque, Richard A. Bryce
AI summary
Overview
Research area: AI for drug discovery — specifically LLM-based agentic workflow orchestration for early-stage computational drug discovery.
Technical level: Intermediate. Readers should be comfortable with language-model fine-tuning concepts (LoRA, tool calling, JSON schemas) and general cheminformatics terminology (ADMET, SMILES, pIC50, QED).
Scope: The paper presents a locally deployable, compact-model agentic "Planner" that turns natural-language drug-discovery requests into valid multi-tool execution plans across 18 modular tools, trained on 1,263 workflow-graph-derived query–plan pairs.
What This Paper Is About
Early-stage computational drug discovery requires chaining many different scientific tools — molecular generation, property prediction, ADMET assessment, docking, synthesis planning — in an order that changes as new evidence arrives. General-purpose LLMs can orchestrate these tools, but repeatedly calling remotely hosted models introduces API costs, provider and network dependence, and the risk of exposing project-level research strategy. This paper asks whether a much smaller language model, fine-tuned and run on local hardware, can reliably perform that orchestration — and finds that it largely can, but struggles to compose entirely new workflow sequences.
Key Contributions
-
A task-oriented organization and modular tool ecosystem for early-stage drug discovery: 18 tools grouped into four functional groups (Molecular Design and Generation, Molecular Developability, Target Engagement, Knowledge Extraction) plus a cross-group Multi-Objective Ranker, all accessed through a standardized plug-in interface so individual tools can be replaced or extended.
-
A path-coverage-based data construction strategy that derives supervision from an explicit workflow graph. Instead of collecting isolated function calls, the authors traverse a directed graph of valid tool dependencies, convert each valid path into a structured plan, and then use GPT-5.5 to generate linguistically diverse queries whose intent matches that plan — yielding 1,263 manually reviewed query–plan pairs (including 467 cross-group workflows).
-
A Unified Molecular Schema (UMS) for structured state management: a molecule-centric, block-structured record (blocks for physchem, affinity, scaffold, pocket, docking, md, admet, toxicity, retrosynthesis, novelty, and artifacts, plus project-level pipeline state) that lets heterogeneous tools read and write only their authorized fields across iterative execution and replanning cycles.
-
A trajectory-level evaluation of compact planners. Rather than judging only the final answer, the paper measures whether the generated execution plan is parseable, schema-compliant, correctly ordered, and correctly parameterized — including under a split that deliberately excludes seen workflow compositions.
Main Findings
-
Fine-tuning makes plan generation essentially reliable on in-distribution-style queries. Under the query-level split, on 47 held-out cross-group queries, all three fine-tuned models (Llama 3.2-3B, Qwen3-4B, Gemma 3-4B) reached parse success of 1.000 and schema success of 1.000.
-
A 3B-parameter model is the strongest planner in the query-level setting. Llama 3.2-3B achieved a tool-selection F1 of 0.998, sequence exact match of 0.979, sequence LCS of 0.996, and argument F1 of 0.960 after fine-tuning, versus zero-shot values of 0.786 tool F1, 0.000 sequence exact match, and 0.658 argument F1 (with parse and schema both 0.553 zero-shot).
-
Zero-shot compact models fail at ordering, not at tool identification. Qwen3-4B reached only 0.298 parse and 0.298 schema success zero-shot with 0.778 tool F1 and 0.000 sequence exact match; Gemma 3-4B reached 0.638 parse, 0.574 schema, 0.871 tool F1, and 0.148 sequence exact match. Fine-tuning brought both to 1.000 parse and 1.000 schema, with sequence exact match of 0.915 (Qwen3-4B) and 0.936 (Gemma 3-4B).
-
Compositional generalization remains the hard problem. Under the stricter workflow-grouped split, which excludes identical ordered tool sequences across partitions, sequence exact match drops to 0.452–0.548, showing that compact models struggle to generate complete workflow paths unseen during training.
-
A CDK2 case study shows the system end-to-end. The Planner selected MolGPTGenerator (producing 20 candidates), PhysChemEvaluator, AffinityPredictor, and the rule-based MultiObjectiveRanker, which weighted target engagement at 70% and physicochemical developability at 30%. Eight of the 20 generated molecules exceeded the predefined potency threshold of predicted pIC50 ≥ 6.0. The rank-1 candidate scored 0.724 overall with predicted pIC50 of 6.578, QED of 0.741, and no Lipinski rule-of-five violations; rank 2 scored 0.684 (pIC50 6.381, QED 0.744), and rank 3 scored 0.651 (pIC50 6.263, QED 0.721).
-
Training setup is lightweight and reproducible. LoRA fine-tuning used rank 16, scaling parameter α = 32, dropout 0.05, maximum sequence length 4,096 tokens, an initial learning rate of 2×10⁻⁴, up to three epochs with early stopping, on one NVIDIA H200-SXM GPU with 141 GB of memory, with an 80:10:10 train/validation/test split of the 1,263 pairs.
-
The paper explicitly frames a gap in prior work: existing agentic drug-discovery systems often rely on remotely hosted general-purpose LLMs and encode workflow knowledge through hand-curated documentation or hard-coded prompt stages rather than a machine-checkable dependency structure. It cites the reported AstraZeneca ChatInvent migration experience, where moving from single-agent to multi-agent architecture improved efficiency but increased tool-call errors needing largely ad hoc prompt engineering.
Methodology in Plain English
The authors start by treating each of their 18 tools as a node in a directed graph. An edge from tool A to tool B means B can validly run after A, based on whether B's required inputs will already exist in the shared molecular record. They then walk this graph to enumerate valid multi-step workflows — mostly top-down through the functional layers, but also cross-layer, intra-layer, and reverse-direction paths that begin from retrieval of known compounds rather than de novo generation. A path is only kept if every required input is available from the user query, the initialized state, or a preceding tool's output.
Each retained path becomes a structured JSON plan with ordered, parameterized tool calls. The authors then prompt GPT-5.5 to write several natural-language user requests whose intent matches that plan — the model generates linguistic variety, not the workflow — and manually review and refine the queries, removing invalid paths and inconsistent examples. The result is 1,263 query–plan pairs.
They fine-tune Llama 3.2-3B, Qwen3-4B, and Gemma 3-4B with LoRA on those pairs. At inference, the fine-tuned model acts as the Planner inside a LangGraph implementation: a user query becomes an executable tool sequence, each tool reads and writes its authorized blocks in the Unified Molecular Schema, and an Evaluator inspects the resulting records — either returning filtered results to the user or feeding errors back to the Planner for replanning until success or a retry limit.
Evaluation is deliberately at the trajectory level, not just the answer level: they check whether plans parse, whether they comply with the schema, whether the right tools are chosen, whether the order is right, and whether arguments are correct — under both a query-level split and the stricter workflow-grouped split.
Why This Matters
Impact on research. The paper makes a case that reliable scientific tool orchestration does not necessarily require frontier-scale, remotely hosted models. It also argues that evaluating agents only on final answers is insufficient for scientific workflows, because an invalid intermediate action can propagate errors through downstream computations even when the final response looks plausible. Its path-coverage data construction offers a reusable recipe for turning a tool dependency graph into planner supervision.
Real-world applications.
- Privacy-sensitive or IP-sensitive drug discovery: running routine orchestration on locally controlled GPU hardware, with UMS block-level access restricting what data leaves the local environment when external services are used.
- High-volume virtual screening: where iterative cycles of generation, evaluation, filtering, and ranking push API costs and rate limits to impractical levels.
- Constrained or intermittent-connectivity environments where dependence on model providers and network connectivity is undesirable.
- Extensible internal platform building: the plug-in interface and UMS let teams swap in stronger or proprietary tools (for example REINVENT 4, Mordred, GraphDTA, ASKCOS, GNINA) without changing the surrounding workflow, though the paper notes these alternative implementations were not evaluated.
Industry relevance. The framing speaks directly to teams building internal agentic platforms — the paper points to the reported AstraZeneca experience where upgrading an LLM forced re-validation of agent orchestration, and positions a small, specialized, locally deployed Planner as a more stable and maintainable alternative for the routine, structurally constrained parts of the workflow.
Future Directions
- Closing the compositional generalization gap. Sequence exact match of 0.452–0.548 under the workflow-grouped split is the paper's central open problem; the authors call for a larger, more diverse dataset covering additional tool combinations, argument configurations, failure conditions, and replanning scenarios.
- Realistic data from practitioners. They argue for close collaboration with drug-discovery researchers so that datasets reflect practical scientific workflows rather than only synthetically enumerated paths, and they note the GPT-5.5-generated queries still need manual screening to reduce unnatural phrasing.
- Replacing rule-based components with lightweight models. Workflow evaluation and result summarization are currently handled by rule-based components; specialized lightweight language models could enable more context-sensitive assessment, but would need new task-specific datasets and raise a multi-model memory and compute overhead challenge.
- End-to-end workflow evaluation. The paper states that the scientific accuracy of molecular predictions is determined by the underlying domain-specific tools and algorithms rather than by the Planner, and that future work should test whether tool selection, execution order, state transfer, and iterative replanning collectively preserve the reliability and utility of the resulting drug-discovery outcomes. It also notes the current LightGBM affinity model may improve with more advanced predictors, and that a pocket-aware molecule generator would likely improve affinity scores.
Target Audience
Researchers and engineers building LLM-based scientific agents, particularly those working on tool orchestration, function calling, and trajectory-level evaluation; computational chemists and cheminformatics teams evaluating whether compact local models can replace API-dependent agent stacks; and industry platform teams assessing locally deployable alternatives to remotely hosted general-purpose LLMs for early-stage drug discovery. It is also relevant to readers interested in synthetic supervision data construction via graph path coverage.
Authors’ abstract
Early-stage computational drug discovery requires coordinating heterogeneous scientific tools across multi-step workflows. We present a lightweight, tool-augmented framework in which a locally deployable compact language model plans calls to 18 modular tools. A Unified Molecular Schema maintains shared molecular records, while a plug-in interface supports tool replacement and extension. We construct 1,263 manually refined query-plan pairs through workflow-graph path coverage and apply LoRA fine-tuning to three compact model families. Under the query-level split, all fine-tuned models generate fully parseable and schema-compliant plans on 47 held-out cross-group queries. Llama 3.2-3B achieves a tool-selection F1 of 0.998, sequence exact match of 0.979, and argument F1 of 0.960. Under the stricter workflow-grouped split, which excludes identical ordered tool sequences across partitions, sequence exact match reaches 0.452 to 0.548, highlighting the remaining difficulty for compact models in generating complete workflow paths unseen during training. These results demonstrate the feasibility of compact, locally deployable planning while identifying compositional generalization as an important direction for further improvement.