Research
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Overview Research area: LLM agent engineering — specifically the co-optimization of agent "harnesses" (system prompts, tool sets, execution hooks, context-management scaffolding) and model weights for

- arXiv
- 2609.09134
- Published
- 2026-09-08
- Authors
- Zhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra, Sougata Chaudhuri, Shilpa Bhagavath, Zeyuan Chen, Ran Xu, Phil Mui, James Zhu, Sitaram Asur
AI summary
Overview
Research area: LLM agent engineering — specifically the co-optimization of agent "harnesses" (system prompts, tool sets, execution hooks, context-management scaffolding) and model weights for domain-specific enterprise tasks.
Technical level: Intermediate to Advanced. The paper assumes familiarity with LoRA fine-tuning, agentic tool-calling loops, supervised fine-tuning on expert trajectories, and failure-mode taxonomies for agents.
Scope: A single empirical study across seven enterprise agentic benchmarks showing that expert-trajectory imitation destroys the fit between a model and a harness evolved for it, and that on-policy expert correction avoids this failure while enabling iterative co-evolution.
What This Paper Is About
An agent's success depends on both its underlying model and the harness wrapped around it, and prior work has shown that automatically searching over harnesses can lift small models to near-frontier accuracy at a fraction of frontier inference cost. This paper asks what happens when you also adapt the model weights — the other major lever — after the harness has been optimized for a specific weak model. The surprising answer is that the obvious move (fine-tune the weak model on a stronger expert's successful trajectories under the evolved harness) actively hurts performance, and the paper diagnoses why and proposes a fix.
Key Contributions
-
Demonstration of upward harness transfer. A harness evolved using only a weak model (Qwen3-Coder-30B-A3B) transfers cleanly to a much stronger expert (Gemini 3.1 Pro Preview), which uses it better than the weak model does — revealing model-side headroom that imitation was supposed to capture.
-
Identification of a harness–model interference effect. Full-trajectory expert imitation under the evolved harness regresses the weak model by 4 to 30 points on all seven benchmarks, while the identical recipe helps under the unevolved baseline harness — isolating the failure to the interaction between imitation and harness evolution, not imitation itself.
-
A mechanistic diagnosis via failure-taxonomy analysis. Imitation does transfer the expert's knowledge (the implicit-knowledge failure bucket shrinks), but it also transfers the expert's planning style, which the weak model cannot execute; planning failures rise from 1.1% to 14.6% of hard failures.
-
An on-policy expert-correction pipeline. A meta-level MLE agent localizes the single failing turn in the weak model's own rollouts, and the expert rewrites only that turn, preserving the model's native planning distribution. This yields +1.7 points on top of the evolved harness with no regression.
Main Findings
-
Harness evolution is a large, transferable lever. Evolving the harness lifts the weak Qwen model from 29.2% to 78.0% mean test success (+48.8 points), and lifts the Gemini expert from 84.4% to 93.6% (+9.2) on the same harness. The expert actively uses nearly every evolved harness edit in 93.6–100% of rollouts.
-
Imitation under the evolved harness backfires. Fine-tuning Qwen on Gemini's successful evolved-harness trajectories drops mean success from 78.0% to 63.1% (−14.9), with regressions ranging from −4.2 (anomaly detection) to −29.9 (payroll auditing) across all seven tasks.
-
The same recipe helps on the baseline harness. Under the unevolved harness, identical imitation raises Qwen from 29.2% to 35.5% (+6.3). The teaching signal is not the problem; its interaction with harness specialization is.
-
The regression reproduces across model families. On Webarena, Gemma-4-26B-A4B rises 46.7% → 55.6% from harness evolution, transfers upward to Gemini (71.1% → 81.1%), then falls to 41.1% after imitation — 14.5 points below its evolved-harness baseline and 5.6 points below its default-harness baseline, despite Gemma being a reasoning model stylistically closer to the expert.
-
The failure is planning drift, not lost knowledge or lost harness use. After imitation the weak model uses the harness more, not less (domain-computation recipe adoption rises from 30.8% to 76.1%), and the implicit-knowledge failure share falls slightly (46.2% → 44.5%). But planning failures jump from 1.1% to 14.6% of hard failures.
-
On-policy correction makes the two levers compose. Correcting only the failing turn in the model's own rollouts raises mean success from 78.0% to 79.7% (+1.7), gains on five of seven tasks, and stays within noise on the two tasks where the harness is already near ceiling. Planning failures stay near the base-model floor (1.1% → 1.8%), while knowledge failures fall further (46.2% → 43.2%) and overall failure rate drops from 28.9% to 26.8%.
-
The correction is cheap. The pipeline builds roughly 500 training rows (about 400 expert-corrected turns plus about 50 self-pass examples prioritized at the capability boundary) and a single LoRA training run completes in under an hour.
Methodology in Plain English
The study uses seven objectively verifiable enterprise tasks (payroll auditing, budget approval, stock alerting, IoT anomaly detection, browser automation, website management, and code refactoring) with the task suite and environments of prior work. A weak executor model — Qwen3-Coder-30B-A3B, or Gemma-4-26B-A4B for replication — runs the tasks while a stronger Gemini 3.1 Pro Preview meta-agent reflects on failures and proposes harness edits (prompts, tools, hooks, context management). Edits are kept only when validation performance improves, producing an evolved harness specialized to the weak model.
For the imitation arm, the team collects the expert's successful trajectories under that evolved harness, converts them to the student's chat and tool-call format, mixes in the student's own passing rollouts, and LoRA fine-tunes (rank 16 or 64, 2 epochs, learning rate 1e-4, sequence length up to 98,304 tokens). For the on-policy correction arm, they instead start from the student's own rollouts, use an automated failure-locus step to find the single turn where each failed rollout goes wrong, and have the expert rewrite only that turn — one sentence of corrected strategy plus the corrected tool call — while leaving surrounding steps intact. They sample three candidate corrections per turn and keep the best under a quality judge. To explain the results, an LLM-as-judge classifier sorts every failed rollout into a six-category adaptation-failure ontology (other, tool-use, instruction-following, implicit knowledge, long-context, planning) and compares the failure composition before and after each training recipe.
Why This Matters
Impact on research: The paper challenges the implicit assumption that harness optimization and weight adaptation are independent, composable levers. It shows they become coupled once a harness is optimized around a particular model's planning behavior, meaning naive stacking of the two techniques can undo the gains of one. This establishes an on-policy principle for agent post-training under specialized scaffolds, and connects the harness-evolution literature to the broader teacher-intervention literature (which has mostly studied fixed interfaces).
Real-world applications:
- Enterprise agent deployment where internal API tool-calling, code refactoring, and long-horizon data auditing must run at small-model cost rather than frontier-model cost.
- Cost-constrained agent products that already invest in prompt/tool/harness engineering and want to know whether adding a LoRA fine-tune will help or hurt.
- Training-data pipelines that currently harvest expert trajectories wholesale; the result suggests localizing supervision to the student's own failure states instead.
- Iterative agent improvement loops where harnesses and checkpoints are updated on alternating cycles.
Industry relevance: The recipe is deliberately lightweight — under an hour per training run, LoRA-scale, no RL infrastructure — which makes it practical for teams with modest GPU budgets. The concrete design principle (after harness evolution, supervise at states the student actually visits, not at states the expert visits) is directly actionable for anyone building domain-specific agents.
Future Directions
- Combine on-policy correction with reinforcement learning to push the weak model further where capability headroom remains, since LoRA-SFT on corrected turns does not fully close the gap to the expert.
- Make harness evolution aware that the model will later be fine-tuned, so the two arms are jointly optimized rather than run sequentially — the current loop still evolves the harness first and adapts the model second.
- Characterize when model–harness fit is worth preserving. The paper shows it matters after harness specialization, but the conditions under which sacrificing fit for knowledge transfer is net positive need a general account.
- Extend beyond LoRA-scale adaptation. Whether the planning-drift effect shrinks with higher-rank or full fine-tuning, or with larger students, is untested.
Target Audience
Researchers and engineers working on LLM agents, agentic post-training, and cost-efficient enterprise deployment. It is most valuable to practitioners who already do harness or prompt optimization and are considering adding a fine-tuning stage, and to researchers studying imitation versus on-policy supervision, teacher–student distillation for agents, or failure-mode analysis of long-horizon tool-using systems. Readers without background in agentic evaluation or parameter-efficient fine-tuning will need to consult the cited prior work for context.
Authors’ abstract
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harness and model weights shape behavior, we ask how harness evolution and lightweight fine-tuning should be combined. Across seven enterprise agent tasks, we first evolve a harness with the weaker model, then find that a stronger expert often uses it more effectively, suggesting expert supervision could close the remaining gap. However, training the weaker model on the expert's complete trajectories under the evolved harness backfires: performance regresses on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, even though the same procedure helps under the unevolved harness. Our analysis shows that imitation transfers knowledge and increases scaffold usage, but disrupts model-harness fit: the weaker model adopts the expert's planning strategy without the competence to execute it and no longer matches the harness evolved around its native planning style. We therefore develop an on-policy expert-correction pipeline, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn. This preserves the model's planning style and combines the gains of harness evolution and model adaptation. Our results identify and resolve a source of contention between harness and weight updates, yielding a compatibility-preserving recipe for economical co-evolution on domain-specific enterprise tasks.