Research
Can AI Scientists Coordinate at Runtime?
Can AI Scientists Coordinate at Runtime? Overview Research area: Multi-agent AI systems for automated scientific research (AI scientists), specifically runtime coordination and next-agent selection. T

- arXiv
- 2610.00980
- Published
- 2026-10-01
- Authors
- Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li, Yu Chen, David Xu, William F. Shen, Xinchi Qiu, Xisen Wang
AI summary
Can AI Scientists Coordinate at Runtime?Overview
Research area: Multi-agent AI systems for automated scientific research (AI scientists), specifically runtime coordination and next-agent selection.
Technical level: Advanced — the paper formalizes coordination as a state-dependent decision policy and evaluates it across three host systems, though the core argument is stated in plain terms.
Scope: A single-seed exploratory study introducing Runtime Agent Coordination (RAC), which lets AI-scientist hosts select which existing agent acts next based on the current research state, tested across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench plus an exploratory DiscoveryBench subset.
What This Paper Is About
Most multi-agent AI scientists coordinate through design-time orchestration: fixed workflows, predefined stages, loops, and fallback rules. Human scientists, by contrast, adjust their division of labor as a project unfolds. The paper asks whether AI scientists can also coordinate at runtime — selecting which authorized agent should act next based on current artifacts, open problems, execution history, and remaining budget — rather than executing a predefined sequence. The goal is to test whether a shared runtime coordination layer improves research performance across hosts with different native degrees of flexibility, under matched budgets.
Key Contributions
-
Formal problem formulation. The paper formulates next-agent selection in multi-agent AI scientists as a state-dependent decision, explicitly distinguishing coordination (which agent acts next, what it should produce, whether its output is valid) from communication (what agents share). It defines a decision step where the next agent is either chosen by the host's native successor policy (N0, R1) or by a shared runtime selection policy (R2, R3), constrained to the agents the host currently permits.
-
A host-preserving instantiation. The RAC framework is instantiated in Agent Laboratory, EvoScientist, and ARK without replacing their native agents, tools, or permissions. Each host bridge serializes native state, invokes an existing capability, and returns artifacts and usage records through a shared interface; bridges contain no host-specific runtime selection rules, verification criteria, or benchmark hints.
-
A cumulative four-condition experimental design. Four conditions separate native execution (N0), runtime communication (R1), runtime selection (R2), and the joint addition of scoped contracts and artifact-grounded verification (R3), with base models, native role prompts, tools, permissions, artifact interfaces, and task inputs held fixed within each host.
-
Empirical reporting of scores, tokens, and costs. The paper reports ResearchClawBench weighted final scores, input and output token usage, and standardized inference-cost estimates across hosts and conditions, plus an exploratory transferability test on five DiscoveryBench tasks.
Main Findings
-
Runtime selection produced the highest observed mean score for every host. R2 scored 18.42 for ARK, 12.08 for Agent Laboratory, and 18.53 for EvoScientist, versus N0 scores of 16.66, 5.47, and 15.99 respectively (Table 3).
-
Adding contracts and verification reduced means relative to R2 in every host. R3 scored 17.98 for ARK, 9.88 for Agent Laboratory, and 15.96 for EvoScientist — each below the corresponding R2 mean, though R2–R3 treats contracts and verification as a joint addition rather than isolating their individual effects.
-
Full RAC (R3) versus native execution was host-dependent. R3 increased mean scores for Agent Laboratory (5.47 to 9.88) and ARK (16.66 to 17.98), while EvoScientist remained nearly unchanged (15.99 to 15.96). Recorded input-token usage and estimated costs decreased for Agent Laboratory but increased for the other two hosts.
-
Communication alone helped two hosts but not the third. R1 scored 17.40 for ARK and 9.63 for Agent Laboratory, both above their N0 baselines, but 12.07 for EvoScientist, below its N0 score of 15.99.
-
Mean gains did not imply uniform benefits. R2 increased scores over N0 by 6.61 points for Agent Laboratory, 2.55 for EvoScientist, and 1.76 for ARK, but both gains and declines occurred within every host and condition. R3 improved Math-000 and Math-001 for both Agent Laboratory and EvoScientist but reduced EvoScientist's scores on Physics-002, Math-002, and Math-003. Gains persisted across all three conditions for ARK on Physics-002 and Agent Laboratory on Life-001.
-
An exploratory DiscoveryBench subset favored R3. Across five exploratory tasks and six task–host pairs (Agent Laboratory covered all five tasks; EvoScientist covered only NLS SES), mean improvements over N0 were 19.05 points under R1, 18.07 under R2, and 25.92 under R3. R3 achieved the largest aggregate gain, but the subset was not randomly sampled and does not establish representative cross-benchmark performance.
-
R2's slightly lower DiscoveryBench mean than R1 traced to two tasks. In the incarceration task, R2 analyzed simulated data and reported synthetic year indices rather than the requested survey years. In archaeology, additional quantities in R2's answer reduced the evaluator's variable-matching score while relation and context matching remained unchanged. Both runs completed the same six-stage sequence as R1 without exhausting their budgets.
-
Trace analysis surfaced three distinct coordination effects, from 1,181 hop excerpts and 180 external scoring receipts across 60 cells. An evidence–action closure effect: in Physics-002 with EvoScientist R2, the reviewer reported at hop 6 that the report contained only 39 bytes; at hop 7 the selected debug capability reported producing a 337-line report of approximately 22 KB, with ten recorded scientific-artifact changes. A planning stagnation effect: in Physics-002 with EvoScientist R3, the reviewer identified a standard-error discrepancy of 0.011 in the report versus 0.015 in the data, but hops 8–13 remained planner calls with no recorded scientific-artifact changes, ending at the 14-hop limit. A self-consistent drift effect: in Neuroscience-000 with EvoScientist R3, the native reviewer passed consistency checks and assigned 8/10, while external evaluation identified missing cross-laboratory, sex, and environment SHAP comparisons and yielded 13.6; separately, Life-001 with Agent Laboratory R2 analyzed ten synthetic patients rather than the supplied seven-patient dataset.
-
Resource patterns differed by host. Agent Laboratory recorded its lowest input-token and cost means in R2 (2.36M input tokens, US$4.33) but its lowest output-token mean in R1 (269.91K). EvoScientist recorded its lowest token and cost means in N0 (11.49M input tokens, US$15.44). ARK recorded its lowest input-token, output-token, and cost means in R2 (45.47M input tokens, 1,113.58K output tokens, US$64.42).
Methodology in Plain English
The researchers took three existing AI-scientist systems — Agent Laboratory (which uses fixed research phases), EvoScientist (which adds persistent memory and self-evolution), and ARK (a human-steered harness) — and wrapped them with a shared coordination package without swapping out their agents, tools, or permissions.
They then ran each host under four cumulative conditions. N0 runs the host as-is. R1 adds a communication channel between agents but keeps the host's native choice of who acts next. R2 keeps that channel and lets the current agent pick the next authorized capability based on a checkpoint containing the objective, saved artifacts, unresolved problems, execution history, and remaining budget. R3 adds two more things: a scoped work contract specifying the objective, readable and writable artifacts, and required output; and artifact-grounded verification that checks persisted artifacts against the contract and records a verdict of supported, refuted, or inconclusive. Critically, every verdict preserves the workspace and host transition — feedback enters the next selected agent's prompt rather than triggering retries, rerouting, rollback, or stopping. Runs end on budget exhaustion, hard hop limits, terminal provider errors, or native host completion.
Evaluation used ResearchClawBench, which supplies real-paper-derived tasks with a research question, related literature, raw data, and an executable environment while withholding the target paper, spanning mathematics, neuroscience, information science, energy, life science, and physics. Agent Laboratory and EvoScientist each covered ten tasks; ARK covered five. All runs used seed 0. DeepSeek V4 Pro served as the execution model, and GPT-5.5 as the external ResearchClawBench judge — separate from the in-loop verifier and unable to influence agent selection or host transitions. Hop ceilings were calibrated per host to its own N0 hop count rather than set uniformly, and coordination and verification work was charged to the run's budget. Cross-run memory and evolution stores were reset before each run to isolate within-run coordination. Costs were estimated from recorded tokens at a uniform DeepSeek V4 Pro peak tariff of US$1.32 per million non-cached input tokens and US$3.96 per million output tokens, with no cache-hit or off-peak discounts and excluding external judging charges.
Why This Matters
Impact on research. The paper reframes a design question that most AI-scientist systems answer in advance: who acts next. Its finding that runtime selection produced the highest observed mean for all three hosts, while contracts and verification reduced means relative to selection, suggests that adding coordination machinery is not automatically beneficial and can consume budget that would otherwise fund research work. The trace analysis also separates internal consistency from coverage of the research objective — a distinction that matters for anyone trusting agent self-reports. The authors are explicit that these are descriptive, single-seed results and that verification is not a safety gate.
Real-world applications:
- Automated research pipelines that must decide, mid-run, whether a missing result calls for an experimenter, a broken bibliography calls for a deterministic check or writer, or a contradictory figure calls for code rather than another prose revision.
- Cost-constrained deployment of multi-agent research systems, where standardized token-based cost estimates inform whether to spend budget on coordination overhead or on additional research work.
- Benchmark design and evaluation tooling, since the work distinguishes in-loop verification feedback from external benchmark evaluation of final outputs.
- Auditing and provenance for AI-generated research artifacts, via the checkpoint, request, result, contract, verification, and provenance records the framework writes.
Industry relevance. Organizations running multi-agent AI systems under fixed budgets can read this as evidence that dynamic agent routing is worth testing before adding verification layers, and that per-host calibration is necessary — EvoScientist, which already has native adaptation, responded differently from Agent Laboratory's fixed phases. The reported cost accounting convention also gives practitioners a concrete template for estimating inference spend. The paper states that the standardized USD figures are token-based estimates, not provider invoices.
Future Directions
-
Budget sweeps with per-hop usage and artifact changes. The authors argue these would help distinguish whether weaker R3 outcomes stem from resource constraints or from ineffective delegation, and note that additional budget might accommodate R3's overhead even though the repeated planning trajectory shows diagnosis need not lead to implementation.
-
Separating contracts from verification. R3 jointly adds scoped contracts and artifact-grounded verification, so their individual effects remain unresolved. Isolating them is a direct open question the paper raises.
-
Testing variability across independent runs. The current evaluation is single-seed and does not measure variability, so the reported rankings are descriptive and do not establish robustness. Bounded benchmark tasks also do not establish performance in open-ended discovery.
-
Explaining host-dependent outcomes. The paper hypothesizes that host structure shapes the trade-offs — for example, that Agent Laboratory's fixed phases may benefit from information carried between stages while EvoScientist already incorporates native adaptation — and that differences in native prompts and token consumption may affect how effectively hosts use additional context. It states these remain hypotheses, since aggregate scores and usage do not establish why individual runs improve or decline.
Target Audience
Researchers and engineers building multi-agent AI systems for scientific or long-horizon task automation, particularly those working on orchestration, agent routing, or verification; benchmark designers interested in evaluation protocols that separate in-loop feedback from external scoring; and practitioners deploying multi-agent research pipelines under hard budgets who need to weigh coordination overhead against useful work. Readers interested in the formal treatment of state-dependent agent selection will find the problem formulation most useful, while those focused on empirical resource trade-offs will find the per-host token and cost tables most directly applicable.
Authors’ abstract
Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime? To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets. Code is available at https://github.com/systemind-team/Runtime-AI-Scientist.