Research
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
Overview Research area: Large language model (LLM) agents, agent skill optimization, and retrieval-augmented generation. Technical level: Intermediate. The paper assumes familiarity with LLM agents, e

- arXiv
- 2609.38024
- Published
- 2026-09-29
- Authors
- Jaewon Chu, Ji Soo Lee, Jihwan Park, Dohwan Ko, Jeehye Na, Seunghun Lee, Taehoon Lee, Minseo Yoon, Minseok Joo, Yunyang Xiong, Hyunwoo J. Kim
AI summary
Overview
Research area: Large language model (LLM) agents, agent skill optimization, and retrieval-augmented generation.
Technical level: Intermediate. The paper assumes familiarity with LLM agents, execution harnesses, prompt/skill optimization, and BM25 retrieval, but the framework itself is described in mostly conceptual terms.
Scope: This paper introduces Retrieval-Augmented Skill Optimization (RASO), a framework that uses an external corpus of publicly shared skills as prior knowledge for both initializing and iteratively updating an agent's natural-language skill.
Note on completeness: the paper content available to me is truncated partway through the SpreadsheetBench description in Appendix A.3, so some benchmark and implementation details (for example, the size of the GitSkills corpus, and the ALFWorld and WebShop split sizes) are not reported in the content I can see.
What This Paper Is About
An agent skill is a reusable natural-language document that tells an agent how to act under a given "harness" — the set of tools, file access, observation format, and scoring rules an agent operates within. Existing skill-optimization methods build and refine these skills almost entirely from the agent's own expensive execution rollouts, ignoring the millions of skills already publicly shared. RASO's goal is to use that accumulated external knowledge as a prior, retrieving relevant procedural knowledge from other skills and adapting it to the target task and harness so that an effective skill can be built with fewer rollouts and refined more effectively once rollouts are available.
Key Contributions
-
RASO, a retrieval-augmented skill optimization framework that treats an external skill corpus as prior knowledge for both skill initialization and skill update, rather than relying only on the optimizer's parametric knowledge or on rollout experience.
-
Cross-Harness Adaptation, a shared operation that rewrites retrieved procedural knowledge into the vocabulary of the target domain and harness — removing domain-specific nouns, dropping procedures with no counterpart in the target harness, and preserving tool/parameter claims only when the harness description corroborates them.
-
Two complementary stages: Retrieval-Augmented Skill Initialization (RASI), which builds a knowledge-grounded initial skill with no agent rollouts, and Retrieval-Augmented Skill Update (RASU), which retrieves missing knowledge guided by execution feedback and failure modes observed in rollouts.
-
Empirical validation across four benchmarks and two models, showing RASI improves performance without rollouts and RASO with RASU outperforms strong retrieval-free skill optimizers, with ablations isolating the contributions of retrieval, adaptation, initialization, and update.
Main Findings
-
Retrieval alone is not enough, and can hurt. With GPT-5.6-Luna, SkillRouter (which retrieves a skill without Cross-Harness Adaptation) scored 11.44 on OfficeQA and 55.97 on ALFWorld, versus 11.44 and 64.43 for the No Skill baseline — underperforming even no skill on several benchmarks. On SpreadsheetBench, 92.9% of retrieved skills come from a different harness, which the paper identifies as the source of this mismatch.
-
RASI outperforms retrieval-free initialization without any rollouts. With GPT-5.6-Luna, RASI beat RFSI by +5.63 on OfficeQA (45.74 vs. 40.11), +4.77 on Spreadsheet (49.17 vs. 44.40), +3.24 on ALFWorld (72.64 vs. 69.40), and +1.17 on WebShop (45.06 vs. 43.89). With Qwen-3.5-9B the gains over RFSI were +5.61, +2.74, +5.72, and +10.73 respectively.
-
RASO outperforms existing skill optimizers on both models. With GPT-5.6-Luna, RASO reached 49.03 on OfficeQA, 63.33 on SpreadsheetBench, 74.13 on ALFWorld, and 46.61 on WebShop. That is +3.49 over the strongest competitor on OfficeQA (49.03 vs. 45.54), +6.31 on SpreadsheetBench (63.33 vs. 57.02), +1.49 on ALFWorld (74.13 vs. 72.64), and +1.07 on WebShop (46.61 vs. 45.54). With Qwen-3.5-9B, RASO reached 42.25, 31.55, 51.00, and 24.73, improving over the strongest competitor by +4.85, +1.79, +7.47, and +11.30.
-
Initialization and update gains are complementary. Relative to the retrieval-free RFSI + RFSU baseline (40.70 on OfficeQA, 51.67 on SpreadsheetBench), swapping in RASU alone gave 47.56 and 61.07; swapping in RASI alone gave 45.93 and 58.45; combining both gave 49.03 and 63.33 — improvements of +8.33 and +11.66 points.
-
Cross-Harness Adaptation is doing real work. Removing adaptation from RASI dropped OfficeQA from 45.74 to 41.86 and SpreadsheetBench from 49.17 to 41.43. Removing it from RASU (under the same RASI initialization) dropped OfficeQA from 49.03 to 46.70 and SpreadsheetBench from 63.33 to 57.86.
-
RASU beats other update methods under identical initialization. With RASI as the fixed starting skill, RASU improved over WikiSkill on OfficeQA by +1.94 (47.09 to 49.03) and over SkillOpt on SpreadsheetBench by +3.09 (60.24 to 63.33). Relative to RASI with no update, RASU added +3.29 and +14.16 points.
-
Retrieval size peaks at K = 5. Performance improved from K = 1 to K = 5 on both OfficeQA and SpreadsheetBench, with K = 10 slightly degrading performance, suggesting redundant or less relevant retrieved content at larger K.
-
Even a tiny corpus slice helps. Retrieving from only 1% of the external corpus already produced a clear improvement over no retrieval (0%), with further gains as corpus size increased and the largest gains on SpreadsheetBench.
-
Qualitative transfer across mismatched domains. In one SpreadsheetBench example, RASI retrieved a section from an urban remote-sensing skill using numpy/raster arrays. After adaptation, the source's rule of skipping empty groups became "treat the output as missing," and the generated code correctly left department B's cell blank. Removing that retrieved section caused an index shift that moved department C's mean into B's cell.
Methodology in Plain English
The researchers fix both the language model and the execution harness, so the only thing being optimized is the text of the skill itself. Skill optimization is defined to include both building the initial skill and refining it.
Shared retrieval and adaptation pipeline. Skill documents from an external corpus (GitSkills) are split into heading-delimited sections so retrieval can target specific procedures rather than whole documents. A query-generation agent writes queries, BM25 retrieves the top-K sections per query (K = 5 by default), and an adaptation agent converts each retrieved set into a short "lesson" conditioned on the task description, the harness description, and a specific requirement the lesson must resolve.
RASI (initialization). Given only the task and harness descriptions, a query agent produces requirement–query pairs covering procedures and harness constraints. Each query retrieves sections, each set of sections becomes an adapted lesson, and a skill-initializer agent synthesizes the initial skill s₀ by placing each lesson into the corresponding execution step. No rollouts are used.
RASU (update). Starting from s₀, the agent runs rollouts on a minibatch of 40 training tasks. A gradient-and-query generator analyzes failed trajectories, and for each failure mode produces both a textual gradient and a new retrieval query. Retrieved sections are adapted into lessons, and a skill-updater agent writes a candidate skill from the current skill plus the gradients and lessons. The candidate is accepted only if it scores better on the validation set; otherwise the current skill is kept. This repeats for a fixed number of iterations (2 epochs over the training split).
Safeguards. To avoid leaking benchmark-specific knowledge, the authors build a per-benchmark blocklist and remove entire matching repositories from the corpus before indexing. Task and harness descriptions are written once by a coding agent that reads the harness code, without access to test instances, and are held fixed across all methods, models, and seeds. RFSI uses the identical pipeline as RASI minus the retrieved lessons, so the RASI–RFSI gap isolates retrieval rather than prompt or decomposition differences.
Evaluation setup. Four benchmarks — OfficeQA, SpreadsheetBench, ALFWorld, WebShop — with two target models, GPT-5.6-Luna and Qwen-3.5-9B, and three random seeds. OfficeQA covers U.S. Treasury Bulletins spanning nearly 100 years and about 89,000 pages, with 246 questions split into 50 training, 24 validation, and 172 test questions. Baselines are No Skill, SkillRouter, and RFSI for initialization, and TextGrad, GEPA, SkillOpt, and WikiSkill for update. Implementation uses a minibatch of 40 tasks per iteration, chunked into groups of 8 for gradient generation, a Qwen rollout temperature of 0.0 and optimizer temperature of 0.7 served via vLLM on a single NVIDIA RTX A6000 GPU, and "low" reasoning effort for GPT-5.6-Luna.
Why This Matters
Impact on research. The paper reframes skill optimization as a retrieval-and-adaptation problem rather than a pure rollout-driven search problem. It provides a way to incorporate a large, heterogeneous public skill corpus into optimization while explicitly handling the domain and harness mismatch that makes naive retrieval counterproductive — a failure mode the paper documents directly (SkillRouter falling below the No Skill baseline, and 92.9% of SpreadsheetBench retrievals originating from a different harness).
Real-world applications:
- Enterprise document analysis, such as numerical reasoning over large corpora of financial or government bulletins, where OfficeQA-style tasks require both document parsing and computation.
- Spreadsheet automation, where agents must translate natural-language instructions into correct formulas and handle edge cases like empty groups.
- Web navigation and purchasing agents, where WebShop-style interaction budgets and observation formats differ substantially from document-based harnesses.
- Embodied and household task agents, where ALFWorld-style procedural knowledge from other environments can be adapted rather than relearned from scratch.
Industry relevance. Because agent skills are plain text, they are inspectable, auditable, and transferable across models without retraining. RASO makes a shared pool of such skills reusable across harnesses, which reduces the rollout cost of producing a working skill, and its ability to improve a smaller model (Qwen-3.5-9B) suggests the approach can close part of the gap to larger models through externalized procedural knowledge rather than model scale.
Future Directions
- Scaling and curating the skill corpus. The paper shows performance improves as the corpus grows from 1% upward, but does not report the size of the full GitSkills corpus or how corpus composition, quality filtering, or deduplication beyond the benchmark blocklist affects results.
- Moving beyond BM25 and fixed K. Retrieval currently uses BM25 with K = 5 selected empirically, and performance degrades at K = 10. Learned or reranked retrievers, or adaptive per-query K, are natural extensions.
- Better handling of harness mismatch. Cross-Harness Adaptation is an LLM-driven rewriting step with three stated principles; how faithfully it preserves correct tool and parameter behavior, and where it fails, is not quantified beyond the aggregate ablation.
- Reducing dependence on validation-based acceptance. RASU accepts a candidate skill only when it improves validation performance, which requires a validation split and repeated rollouts. Whether retrieval-augmented updates can reduce the number of update iterations or rollouts needed is an open question, since the paper fixes training at 2 epochs.
Target Audience
Researchers and practitioners working on LLM agents, prompt and context optimization, and retrieval-augmented generation — particularly those building or maintaining agent skills for tool-use harnesses. It is also relevant to engineers who want to reuse publicly available skill libraries across different environments, and to readers interested in how external knowledge can substitute for expensive agent rollouts.
Authors’ abstract
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbf{Retrieval-Augmented Skill Optimization (RASO)}, a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: \textbf{Retrieval-Augmented Skill Initialization (RASI)} constructs a knowledge-grounded initial skill without requiring agent rollouts, while \textbf{Retrieval-Augmented Skill Update (RASU)} iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.