Research
Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance
Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance Overview Research area: Machine learning / large language model multi-agent systems, specifically tes
- arXiv
- 2610.02396
- Published
- 2026-10-01
- Authors
- Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li
AI summary
Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution InheritanceOverview
Research area: Machine learning / large language model multi-agent systems, specifically test-time (inference-time) evolution of agent workflows and execution reuse.
Technical level: Advanced. The paper assumes familiarity with LLM agents, directed-graph workflow representations, evolutionary search, and caching/incremental-computation concepts.
Scope: The paper proposes and evaluates Inherit-MAS, a test-time evolution loop for LLM-based multi-agent systems that makes "inheritance" explicit at two levels — which parts of a workflow to keep, and which stored execution results may be reused — and measures its accuracy and token cost on WorkBench and HotpotQA FullWiki.
Paper metadata: Songtao Wei, Yi Li, and Bingzhe Li are affiliated with the University of Texas at Dallas; Zhichun Guo is an independent researcher. The paper is arXiv:2610.02396v1 [cs.LG], dated 01 Oct 2026, licensed CC BY 4.0. Code is released at https://github.com/CrazyMint/Inherit-MAS.
What This Paper Is About
Multi-agent systems built from LLMs divide complex tasks among role-specialized agents, but the workflow connecting those agents — their roles, prompts, tools, and communication links — is hard to design correctly in advance, because weaknesses in evidence collection, tool use, or coordination only become visible after an initial execution.
Test-time evolution tries to fix this by revising the workflow using execution feedback, but two problems remain: broad revisions can disturb components that were already working, and re-executing requests that never changed wastes computation. The goal of Inherit-MAS is to use execution feedback to improve task performance while controlling the inference cost of that evolution, by explicitly deciding what to inherit and what to change.
Key Contributions
-
Test-time MAS evolution formulated as explicit workflow inheritance. After initial synthesis, each candidate starts from the latest completed candidate, keeps the nodes that a node-level critique identifies as useful, and changes the kept workflow through one validated edit directed at the diagnosed deficiency.
-
Conservative change localization and request-verified execution inheritance. Eligible model and read-only tool results can be used directly across related candidates, but only when the complete resolved request and execution context match, and without altering the candidate being evaluated.
-
Empirical outperformance on two benchmarks with two worker backbones. On WorkBench and HotpotQA FullWiki, Inherit-MAS achieves higher completion and joint F1 than EvoAgent, EvoMAS, and TacoMAS, while using less inference than EvoMAS and TacoMAS.
-
Quantified cost savings from execution inheritance. Compared with independent reruns with execution inheritance disabled, execution inheritance avoids 29.1% and 34.6% of live worker inference on the two benchmarks and 5.3% and 18.1% of total tokens.
Main Findings
-
WorkBench completion with GPT-4o-mini workers: Inherit-MAS reaches 55.4% overall completion, which the paper reports as 13.8 points above Single ReAct (41.5%) and 20.8 points above TacoMAS (34.6%), the best evolving-MAS baseline. EvoAgent scores 33.8% and EvoMAS 30.8%, so every evolving-MAS baseline trails Single ReAct.
-
WorkBench completion with Qwen3-32B workers: Inherit-MAS scores 46.2%, matching Single ReAct at 46.2% and exceeding the best evolving-MAS baseline by 11.5 points. EvoAgent scores 28.5%, EvoMAS 33.1%, and TacoMAS 34.6%.
-
Per-domain WorkBench results (GPT-4o-mini): Inherit-MAS scores Analytics 60.0, Calendar 75.0, CRM 20.0, Email 45.0, Project management 35.0, and Multi-domain 67.5. With Qwen3-32B workers it scores Analytics 75.0, Calendar 50.0, CRM 20.0, Email 40.0, Project management 20.0, and Multi-domain 52.5.
-
HotpotQA FullWiki joint F1 with GPT-4o-mini workers: Inherit-MAS reaches 49.7%, reported as 3.7 points above Single ReAct (46.0%) and 1.9 points above TacoMAS (47.8%). EvoAgent scores 34.8% and EvoMAS 39.2%. Inherit-MAS has the highest answer F1 (73.3), while its supporting-fact F1 (61.7) stays at Single ReAct's level (61.9); TacoMAS has the best supporting facts (64.8) but weaker answers (66.8).
-
HotpotQA FullWiki joint F1 with Qwen3-32B workers: Inherit-MAS reaches 42.7%, reported as 5.6 points above Single ReAct (37.1%) and 2.8 points above TacoMAS (39.9%). EvoAgent scores 32.3% and EvoMAS 39.4%. Again Inherit-MAS has the highest answer F1 (65.3) while TacoMAS retains the best supporting facts (60.0).
-
The only evolving-MAS system that matches or exceeds Single ReAct in all four benchmark–backbone settings. Across those settings, the margin over other evolving-MAS systems is widest on WorkBench, where completion is scored on the final environment state.
-
Lower inference than EvoMAS and TacoMAS. Under each system's own configuration, EvoMAS and TacoMAS use 16.8× and 22.9× Inherit-MAS's total tokens on WorkBench and 13.7× and 2.1× on HotpotQA.
-
Performance improves across refinement rounds. Scoring the workflow Inherit-MAS would have returned after initial synthesis and after each refinement round (R1 to R4) raises performance at every round — by 5.0 points in total on HotpotQA and 9.2 points on WorkBench, where the first round contributes 6.9 points.
-
Execution inheritance cuts live worker tokens. Full execution spends 2.03M worker tokens on WorkBench and 5.50M on HotpotQA, versus 1.44M and 3.60M with execution inheritance, saving 29.1% and 34.6% of worker tokens. Because meta-model and judge inference stays live in both settings, total token savings are smaller: 5.3% and 18.1%.
-
Savings concentrate in prompt edits. Within the original trajectories, prompt edits inherit 47.3% and 75.0% of worker-token equivalents. Full execution achieves 55.4% completion in the WorkBench rerun, the same as with execution inheritance.
-
Candidate quality is non-monotonic. Raw WorkBench candidate completion across indices 0–4 is 46.2%, 51.5%, 48.5%, 42.3%, and 50.0%; HotpotQA joint F1 is 44.7%, 46.2%, 47.4%, 45.9%, and 44.4%. The paper therefore describes evolution as dependent iterative search rather than deterministic hill climbing.
-
Ranking matters more than taking the last candidate. On WorkBench, initialization, the last generated candidate, and the online ranking achieve 49.2%, 50.0%, and 55.4% completion; on HotpotQA they achieve 44.7%, 44.4%, and 49.7% joint F1.
-
Remaining ranking headroom. An oracle over stored candidates solves 82 WorkBench tasks versus 72 returned, and reaches 56.3% HotpotQA joint F1 versus 49.7% returned. Offline oracle prefixes rise to 63.1% and 56.3%, indicating that the main remaining gap is that the judge does not always identify the best evaluated dependent refinement.
-
Inheritance is uneven across rounds. At indices 1 through 3, the inherited share of worker-token equivalents is between 34.7% and 44.0% on WorkBench and between 44.0% and 63.2% on HotpotQA. At index 4 it falls to 17.6% and 21.0%, because the final-round restart often produces a fresh workflow; that restart's candidate is the returned workflow for 25 WorkBench tasks and 47 HotpotQA questions.
Methodology in Plain English
Setting up the problem. A multi-agent workflow is represented as a directed graph G = (V, E). Each node is an agent or workflow operation with a role, prompt, tool set, decoding configuration, and declared inputs and outputs; edges carry intermediate results. Executing the graph produces an output, an observable trace of messages and tool interactions, and a cost ledger. The task also comes with an interface specifying allowed roles, tools, input/output formats, and the execution environment. The system searches over directed acyclic workflows within a budget of K candidate attempts.
Initial synthesis and judgment. The first attempt has no predecessor, so a meta-model synthesizes a valid initial workflow from the task and its interface, then executes it and stores eligible node results in an initially empty task-local store. A separately prompted judge then scores the candidate and produces a critique. The judge's score has four components in [0,100]: completion likelihood, requirement coverage, artifact grounding, and overall quality; the critique includes keep/fix recommendations. A candidate counts as "completed" once its execution and a valid judge response are recorded.
Workflow inheritance (two stages). In an ordinary refinement round, Select decides which nodes to keep. It may discard several nodes, but only nodes whose individual removal leaves the workflow valid are candidates — the output node and nodes needed to satisfy required input constraints are excluded. The combined discard set is checked for validity, because individually removable nodes need not be jointly removable; if nothing can be removed, selection is skipped. Edit then proposes exactly one change from a finite menu of operations and targets, addressing the diagnosed deficiency. Edits can change a node's prompt, assigned subtask, or tool access; add or remove a node or edge; reorder inputs; or insert a node between connected nodes. The edit is applied only after a deterministic validator confirms the result is valid — cycles, incompatible connections, and unauthorized tools are rejected. The rest of the kept workflow is inherited unchanged, though the one-edit restriction applies to the kept workflow, not to the entire difference between parent and child. Unsuccessful attempts still consume a candidate slot and count toward failure and cost.
Execution inheritance (localize, verify, reuse). After Select and Edit are validated, change localization compares the new candidate with the original incumbent, accounting for both pruning and editing, to find every new node and every retained node whose specification, ordered incoming wiring, or effective execution settings changed. The conservative affected region also includes everything reachable downstream from those nodes. For a node, the "resolved request" bundles its current specification, ordered upstream records, model and decoding settings, allowed tools, execution limits, and an execution-context fingerprint covering runtime, parser, initial environment state, and fixed tool or retrieval service. The lookup key is a SHA-256 hash of that canonically serialized request. A node inherits a stored result only if it is eligible and a matching record passes format, request-key, and content-hash verification; otherwise it executes live and its result is stored if eligible. Language-model nodes without state-changing tools and deterministic read-only operations can be eligible; state-changing operations never are. Malformed or inconsistent records halt execution rather than being treated as a miss. A hit consumes no live worker tokens, and its original usage is recorded separately as an inherited token equivalent.
Ranking and return. After the attempt budget is exhausted, the system returns the output of the highest-ranked completed candidate without re-executing the workflow. The ranking key is (-b, whether the output is valid, overall quality, requirement coverage, artifact grounding, -index), where b counts recorded node execution errors plus one if the final output is invalid under the benchmark adapter. Comparison proceeds left to right, ties favor the earlier candidate, and cost is excluded from the key. If no candidate completed, the task is recorded as a failure.
Experimental protocol. Inherit-MAS uses GPT-4o-mini workers with GPT-5.4-mini for synthesis/refinement and for the separately prompted judge, all at temperature zero. Each trajectory has an initial workflow plus four subsequent candidates, with one Select decision and one validated edit ordinarily applied per refinement round. A second backbone configuration uses Qwen3-32B workers in non-thinking mode for all five systems, with GPT-5.4-mini meta-models and judges (EvoAgent uses Qwen3-32B throughout). For WorkBench, every candidate runs from a fresh environment copy and a deterministic executor applies its proposed write batch. During a run, the meta-model and judge never observe reference answers, supporting facts, or official metrics; failures and invalid outputs stay in the denominator.
Why This Matters
Impact on research. The paper reframes test-time MAS evolution around two separable decisions: the scope of a revision, and whether a previously computed result is still valid after it. That connects evolving-agent research to ideas from demand-driven incremental computation and build systems, where dependencies determine when prior work remains valid. It also supplies an ablation design that measures execution reuse against independent reruns of the same controller, which separates the benefit of reuse from the benefit of running more computation.
Real-world applications.
- Enterprise workflow automation: WorkBench-style workplace tasks over email, calendar, project management, CRM, and analytics databases, where an agent team must produce a correct final environment state rather than just a plausible answer.
- Multi-document research and evidence synthesis: HotpotQA-style question answering that requires assembling an answer together with sentence-level supporting facts across many documents.
- Cost-controlled agent deployment: Any setting where repeated refinement of an agent pipeline would otherwise re-invoke models and read-only tools for requests that have not changed.
- Tool-heavy agent pipelines: Systems that call search indices, retrieval services, or read-only APIs, where verified result reuse reduces live external calls while state-changing actions are still always executed live.
Industry relevance. The paper reports token counts rather than only accuracy, and shows that evolving-MAS baselines used up to 22.9× the tokens of Inherit-MAS under their own configurations. For organizations paying per-token for agent inference, the combination of higher primary-metric scores and lower inference cost is the central practical claim. The conservative rule that state-changing operations are never inherited, and that every candidate for a stateful task starts from a fresh copy of the same initial environment with a designated sink performing all writes, is directly relevant to deployment safety.
Future Directions
-
Better judges and ensemble ranking. The online ranking recovers part but not all of the variation that evolution creates: the oracle over stored candidates reaches 82 WorkBench tasks versus 72 returned, and 56.3% versus 49.7% HotpotQA joint F1. The paper explicitly motivates better calibrated judges and ensemble ranking as a result.
-
Making the final restart round inheritance-friendly. The final-round restart often produces a fresh workflow, which drops the inherited share of worker-token equivalents to 17.6% on WorkBench and 21.0% on HotpotQA, even though that candidate is the returned workflow for 25 WorkBench tasks and 47 HotpotQA questions.
-
Choosing a return policy when later candidates are not uniformly better. Because raw candidate completion is non-monotonic on both benchmarks, an open question is how to select among candidates without relying on an offline oracle.
-
Verifying when a fresh model invocation would reproduce a stored result. The paper treats this as a separate question and notes deterministic-runtime conditions in its appendix, leaving the boundary of safe reuse under non-deterministic runtimes to be established.
Target Audience
Researchers and engineers working on LLM-based multi-agent systems, automatic agent/workflow design, and inference-time compute control will get the most from this paper, particularly those interested in the cost side of test-time evolution rather than accuracy alone. It is also relevant to practitioners building production agent pipelines who need to reduce redundant model and tool calls without changing the behavior of the pipeline being evaluated. Readers looking for an introductory treatment of multi-agent systems should note that the paper assumes prior familiarity with workflow graphs, judge-based evaluation, and caching semantics, and that its experimental details are largely deferred to appendices not fully reproduced in the content above.
Authors’ abstract
Multi-agent systems (MAS) built from large language models coordinate specialized agents to tackle complex tasks, but effective workflows are difficult to design in advance. Test-time evolution refines workflows using execution feedback, yet broad revisions can disturb useful components, while re-executing unchanged requests can incur redundant computation. Inspired by the interplay of inheritance and selection in biological evolution, we introduce Inherit-MAS, which makes inheritance explicit at the workflow and execution levels. A meta-model first synthesizes a workflow of worker agents with declared roles, communication inputs, and tool permissions, and a separately prompted judge scores each executed candidate and diagnoses its deficiencies. In ordinary refinement rounds, \emph{workflow inheritance} starts from the latest completed candidate, may discard removable nodes judged unhelpful, and applies a validated edit to address the diagnosed deficiency. When the new candidate executes, \emph{execution inheritance} inherits eligible stored results only if the complete resolved request and execution context match, avoiding redundant model and tool calls. With GPT-4o-mini workers, Inherit-MAS achieves 55.4\% completion on WorkBench and 49.7\% joint F1 on HotpotQA FullWiki, outperforming EvoAgent, EvoMAS, and TacoMAS. With Qwen3-32B workers, it also exceeds these evolving-MAS baselines on both benchmarks. Compared with rerunning the same controller with execution inheritance disabled, execution inheritance reduces worker-token usage by 29.1\% on WorkBench and 34.6\% on HotpotQA, and total token usage by 5.3\% and 18.1\%.