Research
Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction
Overview Research area: LLM-based multi-agent systems, automated agent/workflow design, and search-guided decision making (cs.LG). Technical level: Advanced. The paper assumes familiarity with reinfor

- arXiv
- 2610.04137
- Published
- 2026-10-02
- Authors
- Som Sagar, Shasha Li, Hejie Cui, Ransalu Senanayake, Sercan Ö. Arık
AI summary
Overview
Research area: LLM-based multi-agent systems, automated agent/workflow design, and search-guided decision making (cs.LG).
Technical level: Advanced. The paper assumes familiarity with reinforcement learning, Monte Carlo tree search, policy/value networks, LoRA fine-tuning, and LLM agent frameworks.
Scope: The paper introduces SHIFT, a method that builds a custom multi-agent "harness" (roles, instructions, tools, and communication structure) for each individual query by searching over predicted outcomes rather than executing candidate designs, and evaluates it on 9,193 tasks across six benchmarks.
What This Paper Is About
The right agent setup depends on the question: a simple lookup needs no planner or verifier, while reconciling a budget workbook needs a planner, a solver, and a verifier. Deciding which agents, instructions, and tools to use normally requires either running many alternatives at inference time (expensive) or hand-designing workflows (slow and inflexible). SHIFT learns to predict which harness will work and what it will cost, so that Monte Carlo tree search can assemble a per-query harness using only local model inference, executing just the single selected design.
Key Contributions
- Per-query harness construction as tree search. The authors formulate harness building as search over agent, instruction, and tool actions, guided by a learned architect that predicts harness utility without executing the candidate.
- State-of-the-art accuracy against 17 shared-executor baselines. On six benchmarks totaling 9,193 tasks, SHIFT-search attains the highest mean accuracy (79.9%), and SHIFT-value exceeds every baseline in mean accuracy (74.5%) while using 32% fewer execution tokens than the strongest baseline.
- Joint optimization beats partial optimization. Choosing structure, instructions, and tools jointly outperforms choosing only instructions or only tools by up to 9.1 percentage points.
- The learned architect selects accurate, low-cost harnesses and transfers. Value-based selection picks more accurate harnesses at lower execution cost from identical candidate pools, and an architect trained only on GSM8K transfers to MATH-500 without retraining.
Main Findings
- Highest mean accuracy across six benchmarks: SHIFT-search reaches 79.9% mean accuracy, exceeding the strongest baseline (Trace, 72.7%) by 7.2 percentage points; the paired task-bootstrap 95% confidence interval runs from 4.42% to 9.88%.
- Gains concentrate on long-horizon, tool-intensive tasks: On OfficeQA, every baseline stays below 30%, while SHIFT-search reaches 59.3% and SHIFT-value reaches 39.0%; on GAIA, SHIFT-search reaches 68.6% and SHIFT-value 64.7%, exceeding Trace by 30.0 and 15.0 percentage points respectively.
- Statistically supported per-question wins: Against Trace, SHIFT wins on 24 OfficeQA questions and loses on four (p = 2×10⁻⁴, paired sign test), and wins on 19 GAIA questions and loses on seven (p = 0.03).
- A cheaper operating point exists: SHIFT-value attains 74.5% mean accuracy, above every baseline, using 32% fewer execution tokens than Trace and roughly one third those of SHIFT-search (relative cost 27.7 versus 84.3).
- Search cuts cost without losing accuracy: Search matches greedy construction in mean accuracy (79.9% versus 79.8%) while using about 30% fewer execution tokens, and 27% fewer than policy-veto. Value-greedy is cheapest (cost 5.3) but falls to 60.2% accuracy.
- Cost savings are not just fewer tool calls: On OfficeQA, tokens fall by 46% while tool calls barely change, and search is cheaper on 60–98% of questions.
- Joint action space matters: On 500 HotpotQA questions with three runs each, full search reaches 69.4%, versus 67.9% when the architect may only delegate tools and 60.3% when it may only add instructions.
- Learned value selects better harnesses: Value selection raises accuracy by 10.5% on HotpotQA and 2.9% on GSM8K with half and one fifth of the execution tokens; its top pick beats the second-ranked candidate by 11.9% on HotpotQA, and the benefit grows with more candidates.
- Structural heuristics fail: Always picking the most agents, tools, or directives is worse than random selection; picking the fewest agents helps on HotpotQA but collapses on OfficeQA (9.4% versus 59.3% with four or more agents).
- Structure's value is query-dependent: On financial-document questions, harnesses with four or more agents have about 50 percentage points higher accuracy than single-agent harnesses; on arithmetic word problems, they use about 19 times as many execution tokens with little accuracy difference. In training, adding agents or tools raises success by 50% and 34% on OfficeQA and GAIA.
- Gains are not solely from OfficeQA: Without it, SHIFT-search still has the highest mean (84.0% versus 81.4% for Trace).
- Transfer to harder problems: An architect trained only on GSM8K, with a Llama 4 Scout executor, reaches 83.6% on the 500 MATH-500 problems versus 81.2% for a fixed single Coder; on levels 4–5, SHIFT reaches 75.6% versus 71.0%.
- Construction is cheap: Building a harness takes under a second on one GPU, a small fraction of the time needed to execute it.
Methodology in Plain English
SHIFT represents a harness as a directed graph: agents are nodes with a role, a list of directives, and a set of permitted tools, and edges pass outputs between agents (including feedback edges for revision). Every intermediate graph is runnable, which lets the method compare partial and complete designs. Construction starts from a minimal harness — a single Coder with no tools — and grows through a vocabulary of 40 actions: 10 structural actions (add a Planner, Researcher, Reasoner, Critic, Verifier, or Aggregator; add feedback loops; grant the full tool set; stop), 15 directives that append instructions to an existing agent, and 15 tool grants. The paper reports this yields about 1.5×10⁷ distinct harnesses.
A small local model, the architect, reads the query, available tools, feasible actions, and a text rendering of the current graph, and outputs two things from its final-token hidden state: a probability distribution over feasible actions and a value estimate of harness utility on a 0–1 scale. Only LoRA adapters and the two output heads are trained; the backbone stays frozen. Monte Carlo tree search (64 simulations, exploration constant 2.5, top 10 actions expanded per node, maximum 10 actions) uses these predictions to explore combinations, so the executor is never called inside the search loop. Search selects the most-visited path's end; SHIFT-value instead executes the highest-predicted-utility candidate from recovered candidates (5 ranked), re-searching with a larger budget (192 simulations) if the value falls below 0.55.
Training uses a cost-aware reward that credits task success and subtracts penalties for execution tokens, latency, timeouts, and additional tools and agents. The policy is trained to match the search visit distribution, and the value head to predict measured utility, with an added ranking loss that trains the value head to preserve the measured utility ordering of harnesses run on the same query.
Two small details in the paper text differ on the architect backbone: the Figure 2 caption names a "local Gemma-2-2B architect," while Section 3.2 describes "a Gemma 4 E2B backbone."
Why This Matters
Impact on research. The paper moves execution out of the per-query search loop, replacing executor calls with local inference from a small learned model. It also shows that agent structure, instructions, and tools should be treated as interacting design choices rather than optimized separately, and that a learned utility predictor can guide search across all three at once.
Real-world applications (drawn from the paper's task domains):
- Document-grounded question answering, such as retrieving a specific number stated in a supplied report.
- Spreadsheet manipulation, such as reading several sheets, performing calculations, editing cells, and verifying the corrected file.
- Multi-hop question answering that requires combining information across sources.
- General-purpose assistance tasks that mix tool use with multi-step reasoning.
Industry relevance. Systems that deploy LLM agents pay per query in tokens and latency, so a method that reaches higher accuracy at 32% fewer execution tokens than the strongest baseline speaks directly to serving cost. The finding that harness construction takes under a second on one GPU while execution dominates runtime means the adaptation overhead can be small relative to the work being done. The approach also fits organizations that already maintain a library of agent roles, instructions, and tools and need a principled way to decide per request which of them to use.
Future Directions
- A single architect that generalizes across datasets and executors. The paper notes that training separate architects per benchmark incurs an upfront compute cost, and that reducing repeated training is future work.
- Better value estimation. The authors suggest improving value estimation using outcomes averaged over multiple executions, since a single execution is a noisy signal.
- More accurate utility balancing. The paper's reward already charges separately for tokens, latency, timeouts, agents, and tools; refining how these are combined across domains is an open area.
- Value beyond one-step lookahead. The failure of value-greedy decoding (60.2%) shows that one-step lookahead cannot credit an agent whose benefit only appears after later actions, leaving room for better ways to attribute delayed value.
Target Audience
Researchers and engineers working on LLM agent orchestration, automated workflow and prompt optimization, or applying search and reinforcement learning to language-model pipelines. It is also relevant to practitioners who deploy multi-agent systems in production and care about the accuracy versus execution-cost trade-off, and to readers interested in agents for document, spreadsheet, and general-assistant tasks. The paper is written at an advanced level, with detailed algorithmic sections and appendices covering the action space, search settings, reward design, and training procedure.
Authors’ abstract
Agent harnesses specify the roles, instructions, tools, and communication structure used to solve a task, and the right harness depends on the query. Because the value of each design choice is observable only through execution, tailoring a harness to each query has required either executing alternatives at inference time or costly manual design. We introduce SHIFT, which moves execution out of the per-query search loop. A local LLM architect learns a policy over harness-building actions from search, and a value function that predicts, from measured executions, a utility balancing accuracy against execution cost. For each query, Monte Carlo tree search uses these predictions to construct a harness. Across 9,193 tasks in six benchmarks, from math to document and general-assistant tasks, with a Gemini 3.5 Flash executor, SHIFT attains the highest mean accuracy, about 80%, outperforming 17 baselines that span prompting, prompt optimization, and workflow search, and exceeding the strongest baseline by 7.2 percentage points. A cheaper mode of SHIFT also attains a higher mean accuracy than every baseline while using 32% fewer execution tokens than the strongest baseline. We further show that choosing structure, instructions, and tools jointly beats choosing only instructions or only tools by up to 9.1 percentage points, and that learned value selection identifies more accurate harnesses with lower execution cost from candidate pools.