Research
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution Overview Research area: Machine learning / LLM agents — specifically automated harness discovery (AHD) for agentic systems, com

- arXiv
- 2609.38349
- Published
- 2026-09-29
- Authors
- Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
AI summary
MILO: Automated Harness Discovery via Orchestrated Multi-Agent EvolutionOverview
Research area: Machine learning / LLM agents — specifically automated harness discovery (AHD) for agentic systems, combining large language model-driven evolutionary search with multi-agent orchestration.
Technical level: Advanced. The paper formalizes harness engineering as strategy-level evolutionary search, defines agents, evaluators, and fitness vectors in formal notation, and assumes familiarity with LLM-based evolutionary search methods (FunSearch, AlphaEvolve, OpenEvolve, GEPA, EvoX, Meta-Harness, Self-Harness).
Scope in one sentence: The paper introduces MILO, a meta-evolutionary framework that co-evolves an agent's harness and its own search strategy — through island-based lineage memory, per-island mutator agents, and an orchestrator agent — and reports state-of-the-art harness discovery results on three long-horizon software engineering benchmarks plus new bounds on open mathematical problems.
What This Paper Is About
Modern AI agents pair a model with a harness, the software layer that governs prompts, tools, context, memory, and control flow (spawning sub-agents, intercepting tool calls, verifying outputs, deciding when to stop). The harness strongly affects performance: the paper cites GPT-5 solving 35.2% of Terminal-Bench tasks with Terminus 2 but 49.6% with Codex while consuming 35% fewer tokens. Building good harnesses is currently artisanal, manual, and expert-driven, and because harnesses are model-specific, the effort must be repeated whenever models change.
The goal is to automate this: given an agent with a fixed (black-box) model, discover a new harness that advances the agent's accuracy–cost frontier over a task distribution. The paper argues existing automated methods explore the combinatorial design space poorly — most optimize only prompts or skills, most discard failed candidates, most fix their search strategy in advance, and nearly all optimize accuracy alone.
Key Contributions
-
A formulation of AHD as diversity-preserving, prune-inducing, exploratory search. The paper formalizes harness engineering as strategy-level discovery of a solver that generalizes over tasks, and proposes MILO, whose memory is partitioned into islands (lineages) that share ideas but never compete for survival, and which is append-only — retaining rejected candidates so mutators can avoid failed patterns and attempt bolder structural changes.
-
A meta-evolutionary orchestrator agent for self-adaptive search. MILO introduces an orchestrator into evolutionary search that detects and escapes local optima by evolving the search strategy itself — mutators, memory, and curriculum. The authors state that, to their knowledge, no prior evolutionary search method adapts as broadly.
-
State-of-the-art results on software engineering and open mathematical problems. MILO outperforms eight state-of-the-art harnesses and six evolutionary search methods across Terminal-Bench 2.1, PaperBench, and DeepSWE, using both a frontier backbone (Opus 4.8) and an open-weight backbone (gpt-oss-120b). It also sets new records on three open problems from EinsteinArena.
-
Evidence that discovered harnesses generalize. The Terminal-Bench 2.1 harness (with gpt-oss-120b) scores 2.7× Mini-SWE-Agent on the harder Frontier-Bench without any further search.
Main Findings
-
Terminal-Bench 2.1 leaderboard-level result: With Opus 4.8, MILO reaches 86.1 ± 2.0%, above the official leaderboard's top entry of 83.8 ± 2.3%, while consuming 26% fewer tokens than its initial harness.
-
Largest gains over prior search methods on all three benchmarks: Measured against MILO's own initial harness with Opus 4.8, resolution rate improves by +12.0% (Terminal-Bench 2.1), +28.3% (PaperBench), and +10.3% (DeepSWE). The best prior search gains are +4.5% (GEPA), +18.3% (Meta-Harness), and 0% (no improvement observed), respectively.
-
Large gains with an open-weight backbone: With gpt-oss-120b, MILO reaches 15.6% on DeepSWE — 20× its initial harness — while prior search stays below 1%.
-
Cross-benchmark generalization of a discovered harness: The harness discovered on Terminal-Bench 2.1 scores 2.7× Mini-SWE-Agent on Frontier-Bench without further search.
-
New bounds on open mathematical problems on EinsteinArena: MILO-evolved harnesses tighten the best-known upper bounds for Erdős minimum-overlap (0.3808586 → 0.3808568) and for the first and third autocorrelation inequalities (1.50274365 → 1.50274360 and 1.45081 → 1.44889). The paper states these surpass the best-known bounds, including those of AlphaEvolve, TTT-Discover, and EvoX.
-
Weak seeds are not fatal to the search: The authors report that the island design lets even the weakest seed yield the best harness.
-
A weak seed is not the whole story for isolation: Islands exchange successful ideas but never compete for survival, confining each mutator's exploitative bias to its own island — a mechanism the authors credit for preserving diversity.
-
Harness choice can equal model choice in impact: The motivating comparison (GPT-5 at 35.2% with Terminus 2 vs. 49.6% with Codex, 35% fewer tokens) and an example where rewriting a tool description cut completion time by 40% support the claim that harness design is a low-cost, high-impact lever.
Methodology in Plain English
MILO runs two coupled loops. The inner loop evolves harnesses; the outer loop evolves the search strategy itself. An agent is written as a pair of a model and a harness, and the model stays frozen throughout — only the harness source code changes.
The heart of the system is a hierarchical lineage memory. Instead of keeping one best candidate or a single Pareto front, MILO keeps a forest of islands, each of which is a tree. Nodes are harnesses; edges are the source-code patches that produced a child from a parent. A node's label records accuracy, cost, evidence (execution traces and verifier reports), and whether the child was admitted. An edge's label records the edit and its resulting accuracy and cost deltas. New harnesses are appended whether or not they are accepted — this is the prune-inducing part, since failed edits remain visible as negative evidence. Islands are seeded with the most behaviorally dissimilar harnesses available and never compete for survival.
Each round, every island runs five stages: (1) select a parent from the island's admitted set using dominance and behavioral-similarity signals; (2) a mutator agent rewrites the entire harness, receiving the parent's source code, its failure evidence, and a UNIX-like indented rendering of the whole island tree; (3) evaluate the child on the search split with multiple rollouts; (4) admit the child only if it advances the island's Pareto frontier over accuracy, token cost, and latency; (5) check whether the island's best held-out accuracy improved, and increment a stall counter if not. Mutators are coding agents with read, write, edit, search, and bash tools, following a diagnose-first workflow and choosing between targeted refinement and structural redesign.
When an island stalls for a fixed patience threshold, the orchestrator agent wakes up. It first diagnoses the bottleneck as one of four types: a dry mutator (children repeatedly rejected or trivial), an exhausted lineage, a crowded-out niche (a dominant lineage suppressing a minority subtree that solves different tasks), or a whole-population plateau. It then applies up to two interventions: Reassign (swap the island's mutator from the pool), Graft (breed against a donor island's best harness as a co-parent, creating a cross-island edge), Speciate (promote a subtree into a new island with its own mutator), and Curriculum (replace the search task set, favoring high-regret tasks where a few harnesses succeed and most fail). This is what makes the search self-adaptive rather than rule-adaptive.
Fitness is explicitly multi-objective. A harness's fitness vector stacks accuracy with an "economy" term on each cost axis, and dominance is defined with an accuracy tolerance so that near-tied harnesses are compared on cost.
Why This Matters
Impact on research. The paper reframes harness engineering as a machine-learning problem amenable to search rather than manual craftsmanship, and argues it is a step toward agents that improve their own harnesses — recursive self-improvement. It also extends evolutionary search in a direction prior work had not: the search strategy itself (memory, mutators, and curriculum) is evolved, not just the prompts or the parent-selection rule. It further makes the case that failure information and cost axes should be first-class in LLM-driven search, where each evaluation takes hours and a single harness evaluation is expensive.
Real-world applications:
- Automated construction and re-tuning of coding agents for repository-level software engineering tasks, where harnesses are model-specific and must be rebuilt when models change.
- End-to-end machine learning engineering agents that run pipelines autonomously over long horizons.
- Reducing inference cost in deployed agents: MILO's Terminal-Bench 2.1 harness consumed 26% fewer tokens than its initial harness at higher accuracy, and the framework treats latency as an optimization axis.
- Scientific and mathematical discovery workflows, as demonstrated by the tightened bounds on open Erdős minimum-overlap and autocorrelation problems on EinsteinArena.
Industry relevance. Production harnesses such as Claude Code and Codex take teams months to build, and open-source alternatives like mini-SWE-agent and DeepAgents require ongoing manual iteration. Because harnesses are model-specific — models differ in tool use, error modes, and prompt sensitivity — the re-tuning burden is perpetual. AHD offers a way to automate that recurring cost without retraining model weights, making harness quality a cheaper lever than model training.
Future Directions
-
Broadening beyond strategy-level discovery. The paper states that its methodology generalizes to other domains and to instance-level discovery, but the reported results cover harness discovery on three benchmarks and three open math problems; how far the orchestrator's diagnoses and interventions transfer to other evolutionary-search settings is left open.
-
Reducing the cost of each evaluation. The paper notes that scoring one harness takes hours, in contrast to fast oracles that let instance-level methods evaluate millions of candidates. Whether the orchestrator's four intervention types can be extended or learned more aggressively to further improve sample efficiency is unresolved.
-
Generalization across models without new searches. The paper demonstrates that a harness found on Terminal-Bench 2.1 transfers to Frontier-Bench with 2.7× Mini-SWE-Agent, but whether a harness discovered for one backbone (Opus 4.8 or gpt-oss-120b) can be transferred or cheaply adapted to a new model remains an open question relevant to the perpetual re-tuning problem.
-
Scope of the orchestrator's action space. The orchestrator currently composes at most two interventions (l ≤ 2) from four options and is triggered by a fixed patience threshold. Whether richer, longer-horizon interventions or alternative stall signals yield further gains is not addressed.
Target Audience
Researchers and engineers working on LLM agents, agent scaffolding, and harness engineering; practitioners of LLM-driven evolutionary search and automated prompt/program optimization; and applied AI teams at companies deploying long-horizon coding or research agents who need to reduce manual harness maintenance and inference cost. Readers should have prior exposure to evolutionary search and agent architecture, since the paper assumes familiarity with methods such as FunSearch, AlphaEvolve, GEPA, Meta-Harness, and EvoX and uses formal notation for evaluation, dominance, and search configuration.
Authors’ abstract
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).