Research
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Overview Research area: Automated harness optimization for LLM agents and automated curriculum learning (cs.AI). Technical level: Intermediate. Scope: The paper proposes ActiveSaddler, a curriculum-le

- arXiv
- 2610.00906
- Published
- 2026-10-01
- Authors
- Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle
AI summary
Overview
- Research area: Automated harness optimization for LLM agents and automated curriculum learning (cs.AI).
- Technical level: Intermediate.
- Scope: The paper proposes ActiveSaddler, a curriculum-learning method that adaptively chooses which training scenarios generate feedback during offline harness optimization, and evaluates it on GAIA2 and Terminal-Bench 2.0.
What This Paper Is About
Existing methods for automatically improving an LLM agent's "harness" (its prompts, tool interfaces, and control logic) focus on how to update the harness, while keeping the set or ordering of training scenarios fixed before optimization begins. The paper argues this is a missing dimension: as the harness changes, the scenarios worth optimizing on also change. The goal is to automatically adapt the training curriculum so that a limited rollout budget is spent on the failures that still matter, rather than on scenarios with little remaining value.
Key Contributions
- Formulates scenario selection for offline harness optimization as an automated curriculum-learning problem, where training scenarios are selected adaptively as the harness evolves.
- Introduces ActiveSaddler, which models the curriculum as a non-stationary bandit with dynamically instantiated "failure-pattern arms," combined with adaptive arm prioritization and adaptive exploration.
- Shows consistent gains on GAIA2 and Terminal-Bench 2.0 over existing harness optimizers and fixed curricula under the same rollout budget, reporting +4.4 and +7.5 percentage points test Pass@1 over the same harness optimizer using a scenario order fixed before optimization.
- Provides ablations identifying three design principles: failure-pattern arms (rather than category- or scenario-level arms), adaptive arm prioritization, and adaptive exploration.
Main Findings
- Strongest test performance: ActiveSaddler reaches 59.8 ± 1.0 Pass@1 on GAIA2 (test split of 300) and 80.0 ± 2.5 on Terminal-Bench 2.0 (test split of 40), the best of the compared harnesses.
- Gains over fixed curricula: On GAIA2 it improves by +3.9 pp and +4.1 pp over category- and scenario-level difficulty ordering; on TB2 by +9.2 pp and +6.7 pp. The paper states it improves over the same AutoSaddler optimizer with a fixed randomly shuffled scenario order by +4.4 pp (GAIA2) and +7.5 pp (TB2).
- Comparison to other optimizers: On GAIA2, AutoSaddler scores 55.4 ± 1.2 and GEPA and Meta-Harness both 54.2 (GEPA ± 2.2, Meta-Harness ± 1.2), versus a manual default agent at 53.6 ± 1.1. On TB2, AutoSaddler scores 72.5 ± 0.0, Meta-Harness 66.7 ± 5.2, GEPA 65.8 ± 5.2, the manual Terminus 2 harness 64.2 ± 2.9, and the manual Terminus-KIRA harness 69.2 ± 3.8.
- Works with a different optimizer: Applying ActiveSaddler to GEPA improves its average Pass@1 from 54.2% to 57.2% (+3.0 pp).
- Better than non-LLM selection rules: EMA-based failure-persistence scoring and count-based UCB-AIR achieve results 3.1 pp and 4.2 pp below ActiveSaddler, respectively.
- Arm representation matters: Replacing failure-pattern arms with category or scenario arms reduces test Pass@1 by 3.0 pp and 3.6 pp on GAIA2, and by 10.8 pp and 6.7 pp on TB2.
- Prioritization matters: Removing the Arm Prioritizer reduces test Pass@1 by 4.5 pp on GAIA2 and 7.5 pp on TB2.
- Adaptive exploration matters: Replacing the Exploration Controller with a fixed schedule that explores every five iterations reduces test Pass@1 by 4.5 pp on GAIA2 and 8.3 pp on TB2.
- Failure-pattern arms hit live failures more often: They exceed category and scenario arms in sampling scenarios on which the current harness still fails by 16.5 pp and 35.3 pp on GAIA2, and by 29.9 pp and 23.8 pp on TB2.
- Exploration shifts over time: On GAIA2 the Draw rate decreases from 44% in the first half of optimization to 20% in the second half; ActiveSaddler shows 2.0× and 6.9× as many unresolved arms at Pull as at Draw decisions on GAIA2 and TB2, while the fixed schedule gives ratios of 0.8× and 1.3×.
- Cost efficiency: On GAIA2, ActiveSaddler reaches 58.5% development accuracy at $298, while AutoSaddler and category ordering need $1,360 and $673 for the same accuracy, and scenario ordering reaches only 56.9% despite spending $1,698. On TB2, ActiveSaddler reaches 78.9% at $128, while AutoSaddler and scenario ordering need $220 and $363, and category ordering reaches only 73.7% despite spending $548.
- Robustness: In an independent GAIA2 run, ActiveSaddler again outperforms all fixed-order baselines.
Methodology in Plain English
The researchers keep the harness-update optimizer fixed and only change which training scenarios feed it. The curriculum is treated as a multi-armed bandit in which arms are not predefined but created on the fly from diagnosed failures: the Failure-Pattern Extractor abstracts each failed execution into a symptom and then normalizes it against the current arm set, linking related failures into a shared "failure-pattern arm" or creating a new one. Each arm keeps a description, the training scenarios that support it, and the associated execution and diagnostic evidence.
At each iteration an Exploration Controller makes a binary decision — Draw (run scenarios never seen before, to discover new weaknesses) or Pull (revisit a known arm) — using the current harness, the arm set, the unseen pool, and the optimization history. If it pulls, the Arm Prioritizer scores each arm with an LLM over four factors — severity (is the failure still active), fixability (can a harness change address it), breadth (how widely would a fix generalize), and side-effect risk (could it cause regressions) — producing a score in [0,1]. These scores become a softmax selection distribution with temperature τ_sel, and an arm is sampled, after which a batch of its supporting scenarios is executed and passed to the harness optimizer. Optimization outcomes then update both the arm priorities and the set of discovered patterns, so the curriculum co-evolves with the harness.
The implementation adds these three components to AutoSaddler, using the GitHub Copilot SDK (GC-SDK) and a shared pattern registry accessed through a pattern command-line interface. Both task agent and optimizer use gpt-5.5, with medium reasoning effort for the task agent and xhigh for GC-SDK optimizer sessions. Train/dev/test splits are out-of-distribution and disjoint, and each optimized harness is evaluated over three test-time executions reporting mean Pass@1 with standard deviation. To prevent regressions, previously successful scenarios become eligible for exploration again once all training scenarios have been explored and every arm visited; a scenario that now fails is processed normally.
Why This Matters
The paper separates two questions that prior work conflated: how feedback is used to update a harness, and which scenarios generate that feedback. It shows the second question carries substantial performance under a fixed budget, and that an adaptive curriculum transfers across benchmarks, across optimizers (AutoSaddler and GEPA), and is better than simple non-LLM scheduling rules.
Real-world applications:
- Building and continuously maintaining agent harnesses for enterprise assistants where task distributions shift over time.
- Cost control for teams that pay per rollout while tuning agents, since better scenario allocation reduces spend to reach a given accuracy.
- Regression control in deployed agents, since unresolved failures are kept as optimization targets and repaired failures can be re-checked.
- Benchmark-driven agent development on long-horizon or terminal-task suites such as GAIA2 and Terminal-Bench 2.0.
Industry relevance: harnesses are often hand-engineered, and the paper reports that ActiveSaddler's optimized harnesses exceed the manual Terminus 2 and Terminus-KIRA baselines on TB2, suggesting automated curriculum-driven optimization can be more effective than manual engineering.
Future Directions
- Whether the same curriculum principle extends to joint optimization of harnesses and model weights, which the related-work section notes as an active direction.
- How the approach behaves online or at test time, since this work is restricted to offline optimization under a fixed train/dev/test protocol.
- How to reduce the additional optimizer-side LLM overhead while keeping the cost-accuracy advantage.
- Whether richer or learned arm representations beyond failure patterns, category arms, and scenario arms could further improve target construction and prioritization.
Target Audience
Researchers and engineers working on LLM agents, agent harnesses, prompt/tool optimization, and automated curriculum learning. It is also relevant to practitioners who must tune agents under limited rollout budgets and to those studying adaptive task selection in reinforcement learning or data-efficient training.
Note: the paper states that a project website and code will be available at https://aka.ms/ActiveSaddler-website. The authors include an AI Use Statement disclosing generative AI tools used for writing assistance and literature retrieval, with all content reviewed by the authors.
Authors’ abstract
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.