Skip to content
AI.info

Research

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Overview Research area: Test-time AI-for-AI (AI4AI) — automated design of execution environments, or "harnesses," for AI agents — sitting at the intersection of agent scaffolding, meta-level learning,

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
arXiv
2609.38143
Published
2026-09-29
Authors
Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang, Heng Ji

AI summary

Overview

Research area: Test-time AI-for-AI (AI4AI) — automated design of execution environments, or "harnesses," for AI agents — sitting at the intersection of agent scaffolding, meta-level learning, and skill transfer.

Technical level: Advanced. The paper assumes familiarity with LLM agent harnesses, execution feedback loops, retrieval methods such as BM25, and ablations with bootstrap confidence intervals.

Scope (one sentence): The paper studies how a Builder model can learn reusable "meta-skills" from a Target model's execution feedback on a development set, then use the frozen skill bank to construct better task-specific harnesses for held-out tasks, with both models' weights remaining fixed.

What This Paper Is About

An AI agent's performance depends not only on its reasoning ability but on the environment in which it acts. This paper asks whether an AI system can learn to build better environments for another AI system — or for itself — without changing either model's weights. The authors separate the system into a Builder, which designs support, and a Target, which uses that support to solve tasks, and they learn reusable support principles called meta-skills from the Target's execution outcomes.

Key Contributions

  1. The meta-skill abstraction. A meta-skill is formalized as a triple (when, provide, use): it identifies observable conditions that call for support, specifies the capability or resource the environment should supply, and explains how the Target should employ that support while retaining responsibility for judgment.

  2. A construction–execution–reflection learning loop. Starting from an empty skill bank, the Builder constructs harnesses for development tasks, reviews Target execution records and scores, and then keeps, revises, or adds at most one evidence-grounded meta-skill per task batch. The bank is frozen after two development-set passes and used to build fresh harnesses for each test task.

  3. A controlled evaluation isolating the value of enactment versus knowledge. Across Harness-Bench and NewtonBench with GPT-5.6-Sol as Builder and three Targets, the full-bank meta-skill Builder reaches a 65.31% macro-average score, 8.95 percentage points above a no-skill Builder with identical construction capability and 12.02 points above delivering the same bank directly to the Target.

  4. Evidence that harness design is a route to system-level self-improvement. When the same model serves as both Builder and Target in three settings, meta-skills improve scores by an average of 18.71 points over no-skill construction and 14.14 points over skills delivered directly to the Target — with no weight updates and no stronger external teacher.

Main Findings

  • Meta-skills work best when the Builder can enact them. Giving the full bank to the Builder beat giving it directly to the Target in all six model–benchmark settings, with per-pair improvements of up to 25.43 points and 12.02 points on average. The Builder converts declarative advice into persistent state, executable tools, verification logic, and control decisions.

  • Experience adds value beyond raw construction capability. Holding the construction space and budget fixed, the full-bank Builder outperformed the no-skill Builder in all six settings by an average of 8.95 points. It also beat full-bank independently learned Target skills by 10.93 points.

  • Gains are largest where coordination is the recurring bottleneck. Against the no-skill Builder, NewtonBench gains averaged 10.96 points across all three Targets, versus 6.95 points on Harness-Bench. NewtonBench requires coordinating experimentation, reasoning, and valid symbolic submission; Harness-Bench spans more heterogeneous workflows.

  • Full-bank access usually beats sparse retrieval. The full skill bank outperformed BM25 top-2 retrieval in five of six settings, with an average gain of 7.19 points on NewtonBench. The single exception indicates broader access is not uniformly beneficial.

  • Skill refinement is non-monotonic. On NewtonBench, the first development pass changed each Target's score by less than 1.1 points, while the second added 13.01 points for Gemini and 12.33 for Qwen; the average gain from zero to two passes was 10.42 points. On Harness-Bench, Qwen dropped 4.06 points from its first-pass peak, suggesting revisions can become overly specific to recent evidence.

  • Transferred meta-skills are useful but implementation-dependent. On NewtonBench, the Sol-to-Qwen bank produced a 13.36-point gain with Sol but a 4.45-point decline when transferred to Gemini-Pro; on Harness-Bench the same transfer produced a 2.72-point gain. All reported confidence intervals for transfer included zero.

  • Recipient-side learning outperformed imported banks. Gemini-Pro's own bank beat the imported Sol-to-Qwen bank by 3.69 points on Harness-Bench and 9.93 points on NewtonBench. Transferring a bank across both Builder and Target (Sol's Gemini-derived bank used by Gemini-Pro for Qwen) still yielded gains of 4.58 points on Harness-Bench and 5.48 points on NewtonBench.

  • Controllers drove most of the measurable harness-component effect. On NewtonBench, removing the execution controller lowered Gemini-3.6-Flash by 13.36 points (95% CI [-18.49, -8.22]), Qwen-Flash by 4.79 points (CI [-10.27, 0.68]), and GPT-OSS-120B by 1.03 points (CI [-6.51, 4.11]). Memory-plus-context removal produced smaller and uncertain effects: -2.05 for Gemini (CI [-7.53, 3.08]), -4.11 for Qwen (CI [-9.25, 1.03]), and +3.42 for GPT-OSS (CI [-2.05, 8.90]).

  • Meta-skills help close the execution gap on NewtonBench. Cases with no valid submission fell from 394 to 254 (35.5%), correct discoveries rose from 439 to 535 (21.9%), with 145 recoveries from no valid submission to a correct law versus 46 regressions. Valid-but-incorrect submissions rose from 43 to 87, showing completion enables but does not guarantee correctness. On Harness-Bench, full successes increased from 59 to 67.

  • Category-level effects vary. All three Targets improved overall on NewtonBench by 6.85 to 13.36 points, with positive estimates for gravity, radioactive decay, and Hooke's law. GPT-OSS-120B improved 5.20 points overall but declined 2.23 points on vertical workflows.

Methodology in Plain English

The authors set up a two-role system. A Builder model (GPT-5.6-Sol in the main experiments) designs an execution environment — a harness — for a Target model (Gemini-3.6-Flash, Qwen3.8-Flash, or GPT-OSS-120B) to work inside. The harness can modify any of seven component families: instructions, memory, context organization, composed tools, execution control, verification and recovery, and workspace preparation. The underlying benchmark tools, scoring rules, and execution budget stay fixed, so the Builder is only reshaping the surroundings, not the test itself.

Skill learning happens on a development split — 11 of Harness-Bench's 106 tasks and 32 of NewtonBench's 324 tasks, roughly 10% in each case. The Builder starts with an empty bank, builds a harness for a development task, observes the Target's public execution record and benchmark score, and reflects on whether to keep, revise, or add a meta-skill. Each batch permits at most one addition or revision, and it must cite supporting evidence from that batch. This runs over two complete development-set passes.

Test-time construction freezes the bank. For each held-out test task, the Builder receives either the full bank or at most two skills selected by a fixed BM25 retriever with positive relevance scores, then constructs a fresh task-specific harness in a clean environment. The Target executes within it and submits.

Evaluation uses deterministic completion-oracle scores on 95 Harness-Bench test tasks and symbolic-structure accuracy on 292 NewtonBench test tasks. Audited harness-construction failures receive a score of zero. All models run at temperature zero with high reasoning effort, a 16K output-token limit for harness generation and an 8K limit for skill update. Execution budgets are 30 turns, 30 tool calls, and 96K cumulative tokens per task on Harness-Bench, and 12 turns, 10 tool calls, and 192K tokens on NewtonBench.

Comparisons include a native environment, a no-skill Builder (same construction space and budget, empty bank), direct delivery of the Builder's bank to the Target, independently learned structured Target skills, and mined free-form Target notes. The paper also runs same-model self-improvement experiments (Gemini-3.6-Flash and Gemini-3.1-Pro serving as both roles), cross-Builder transfer experiments, frozen-harness component ablations, and paired outcome-transition analyses.

Why This Matters

Impact on research. The paper reframes part of agent improvement as environment design rather than model improvement. Its central comparison — same knowledge, different recipient — is a clean control that separates "knowing what help to give" from "being able to build that help." The same-model results suggest a self-improvement axis that operates without weight updates or a stronger teacher, which is relevant to any work on agent scaffolding, automated harness search, and meta-level skill learning. The non-monotonic refinement curves also provide concrete evidence that more reflection is not automatically better, motivating confidence-aware update rules.

Real-world applications:

  • Agent workflow automation: Harness-Bench-style categories such as SRE/release tasks, where a Builder could generate scaffolding — templates, validator scripts, submission checks — for a deployed agent performing operational workflows.
  • Scientific discovery pipelines: NewtonBench-style interactive law discovery, where controllers sequencing experiments, tracking state, and validating symbolic submissions proved most valuable.
  • Tool-augmented assistants: Building persistent memory stores, call sequences, and versioned revalidation into assistants that must maintain consistency across long multi-step tasks.
  • Self-improving deployed systems: Systems that log their own execution outcomes and learn to equip themselves better for recurring task families without retraining weights.

Industry relevance. Teams shipping LLM agents often face a "harness engineering" bottleneck: the model is fixed, but the surrounding prompts, tools, memory, and control loop are bespoke per deployment. This work suggests that the design of that surrounding layer can itself be learned from execution feedback and partially transferred across model combinations — though the transfer results, with confidence intervals that all include zero, caution against assuming it will just work out of the box.

Future Directions

  1. Broader generalization. The authors note that establishing wider applicability requires more diverse Builders and tasks, and that cross-dataset reuse beyond the studied benchmarks is left to future work.

  2. Cost-aware evaluation. The paper does not measure performance gains relative to construction cost; the budget C_x covers only Target execution and explicitly excludes Builder computation for harness construction, and the authors leave cost-relative evaluation to future work.

  3. Smarter skill selection. Because full-bank access beat top-2 BM25 retrieval in five of six settings but not all, the authors call for retrieval methods that account for both skill complementarity and the Target's likely response to support.

  4. Selective retention and rollback. The non-monotonic refinement results motivate an update rule that uses development-only evidence to assess confidence in a revision and decide whether to continue, retain an earlier version, or roll back.

  5. Refining transferred guidance with recipient feedback. Since transferred banks depend on how the receiving Builder implements them, the authors suggest that refining the resulting harness through execution feedback could help realize the value of transferred meta-skills.

Target Audience

Researchers and engineers working on LLM agents, automated agent design, and test-time adaptation will get the most from this paper — particularly those interested in harness/scaffolding construction, meta-level learning, or self-improving systems. It is also relevant to practitioners who build agent infrastructure and wonder whether the environment layer can be learned rather than hand-tuned. Readers without background in agent loops, retrieval baselines, or bootstrap confidence intervals will find the empirical sections demanding; the conceptual framing around meta-skills and the Builder/Target split is accessible without that background.

Authors’ abstract

Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.

Read the original paper