Research
SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
Overview Research area: Training and evaluating large language model (LLM) agents that use externally provided "skills" (reusable packages of task guidance, factual knowledge, or runnable scripts), wi

- arXiv
- 2609.37539
- Published
- 2026-09-29
- Authors
- Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li
AI summary
Overview
Research area: Training and evaluating large language model (LLM) agents that use externally provided "skills" (reusable packages of task guidance, factual knowledge, or runnable scripts), with an emphasis on automatic synthesis of verifiable training environments.
Technical level: Advanced. The paper assumes familiarity with agent harnesses, supervised finetuning, executable verifiers, and agent benchmarks.
One-sentence scope: SkillGym is an automatic pipeline that turns community-written skills into verifiable, executable task environments, collects supervised finetuning trajectories in them, and shows that training improves skill-use performance across model families, sizes, and held-out skills.
What This Paper Is About
Agent skills are now standard in modern agent harnesses, but how to synthesize reliable training data for using them—and how to train agents to apply them—remains underexplored. Existing skill-use training approaches mostly derive skills from an agent's own experience in a small number of fixed environments, which confines learning to a few task domains. SkillGym reverses the direction: it starts from publicly written skills and builds verifiable environments around them, then trains agents on verified successful trajectories so the learned behavior transfers to new skills and tasks.
Key Contributions
- An automatic skill-to-environment pipeline. SkillGym transforms community-written skills into executable training environments, constructing tasks across four reasoning structures (procedural, abductive, constraint satisfaction, partial order) with a builder-reviewer agent system that refines environments, reference solutions, and outcome verifiers.
- Scale of data construction. It constructs 6.8k tasks and collects 19k interaction trajectories (19k verified successful trajectories for supervised finetuning) across multiple agent harnesses, from a curated pool of 11,897 skills.
- Demonstrated training gains and transfer. Supervised finetuning on the resulting data improves six LLMs from three families ranging from 2B to 122B parameters on four skill-use benchmarks, including on tasks involving skills held out from finetuning.
- Behavioral analysis of what training changes. Training teaches agents to consult the provided skills, raising the rate of reading the relevant skill from 28% to 96%, and the gains hold across reasoning structures, including structures that form a minority of the training data.
Main Findings
-
Broad, consistent improvement after SFT. Training improves the model in 22 of the 24 comparisons, with average gains of 13.8 points on the SkillGym test set, 9.7 on SkillEval, 9.7 on SkillsBench, and 41.2 in Skill-Use-Bench SU. The largest and most uniform gain is in skill use itself: SU rises by 21 to 55 points for every backbone.
-
Small trained models rival far larger ones. The 9B model outperforms Qwen3.5-397B-A17B on the SkillGym test set and on SkillEval; the 27B model exceeds two of its three teachers on SkillEval; and the 9B model matches the SkillEval score reported for SKT without using its data, although the two use different harnesses.
-
Gains scale with model size on human-written tasks. Gains on SkillsBench, whose human-written tasks are the hardest, grow with model size within the Qwen3.5 family, from +4.2 at 4B to +23.5 at 122B; the 2B model is the only one whose SkillsBench score drops (10.8 without SFT to 6.1 with SFT).
-
Transfer to held-out skills. SFT increases success from 40.0% to 57.0% on held-in skills and from 42.5% to 62.0% on held-out skills—17.0 and 19.5 percentage points respectively. Held-out success is higher than held-in success for both models, by 2.5 percentage points for the base model and 5.0 points for SFT.
-
Training makes skill access pay off. Providing skills increases base-model success from 33.8% to 41.3%, a gain of 7.5 percentage points, while the SFT model rises from 43.0% to 59.5%, a gain of 16.5 percentage points. SFT improves success even without skills, by 9.2 points, and on SkillsBench the SFT model scores 9.4 without skills and 22.4 with them.
-
The biggest behavioral change is consultation. On Skill-Use-Bench for Qwen3.5-9B, Trigger rises from 28.0 to 96.0, Compliance from 32.5 to 49.0, and Boundary from 57.7 to 64.0. The base model opens the relevant skill in fewer than a third of tasks, which the paper identifies as the main reason its skill-use score is low.
-
Gains are structure-agnostic, but application lags. SFT improves every annotated reasoning structure by similar amounts (16.9 to 22.1 points on SkillGym and 15.4 to 16.4 on SkillEval), and a regression controlling for task difficulty finds no structure-specific effect. Rule application is the primary structure of only 12% of training tasks, yet SkillEval's rule-application tasks improve by 13.9 points. On SkillsBench, tasks whose skills provide reference documentation gain 28.0 points, but those requiring a domain method, a shipped tool, or modification of an existing system do not improve.
-
Quality review adds signal beyond execution validation. At the same token budget, training on review-approved tasks outperforms training on validated-only tasks on every metric, by 4.7 points on the paper's test set, 6.7 on SkillEval, and 9.3 on Skill-Use-Bench completion.
-
Task structure mix is roughly neutral at this scale. Training on a mix of all four task profiles performs on par with training on procedural tasks alone at the same token budget, within run-to-run variation on SkillEval and Skill-Use-Bench, and slightly lower on SkillsBench. The authors note both ablation subsets are much smaller than the full data, and on SkillEval the ablation models fall below the base model because many rollouts end in repeated reasoning that exhausts the output budget.
-
Tasks are skill-critical but not skill-gated. With Kimi-K3, one of the strongest available models and one of the teachers, success rises from 60.5% to 68.8% with skills. 52 tasks are solved only with the skill and 19 only without it (McNemar exact test, p ≈ 10⁻⁴), while 60.5% of tasks are solved without skills. The benefit is largest for procedural tasks (+16.0), then constraint satisfaction (+10.1), abductive (+4.0), and partial order (+3.0).
Methodology in Plain English
The pipeline has three stages.
-
Collecting and curating skills. Skills are downloaded as single folders, each requiring a
SKILL.mdfile. The authors crawl 184k skill entries from skills.sh (around 9.7k top-ranked skills) and the claude-skill-registry GitHub aggregation, keeping 51k unique skills after deduplication. Each skill is annotated on basic properties, runtime requirements, and quality using rule-based checks plus LLM annotation; a human expert independently annotates 100 sampled skills and agrees with the automatic annotations 94% of the time. Skills are then filtered to those that are valid, non-empty, coherently written English, run without runtime network access or GPUs, avoid destructive actions, and stay within at most 300 files. This yields a curated pool of 11,897 skills. -
Constructing skill-critical tasks. Each task must require applying skill-provided knowledge that is not stated in the instruction, and must be reliably verifiable. Every finalized task is a tuple of an instruction, an executable environment, an initial workspace state, an executable verifier, and a reference solution. Four task-profile prompts encode four reasoning structures, each with a difficulty layer requiring realistic input scales, value-based verification that recomputes expected values from inputs, and at least one consequential step depending on non-obvious skill-specific knowledge. A builder-reviewer agent system first builds environments (Dockerfile plus dependencies, validated by the reviewer building the image and running tests), then authors tasks through three gates: a fit-check gate, a validity gate requiring that the verifier fails on the initial workspace and passes on the reference solution's workspace, and a quality gate where the reviewer checks for information leakage and for checks that are too strict to accept valid alternatives or too weak to reject incorrect ones.
-
Collecting trajectories. Training trajectories are generated by three open-weight teacher models—Kimi-K3, DeepSeek-V4-Flash, and GLM-5.2—using four agent harnesses: MiniSwe-Agent (Bash for all actions), AgentFly and OpenCode (dedicated file-operation and command-execution tools), and Terminus-2 (JSON-based command protocol). Diversity is further introduced by varying system prompts, tool names, and tool schemas. A trajectory consists of observations and actions within a task environment.
Training and evaluation. SFT uses the collected trajectories for 2 epochs, a learning rate of 10⁻⁵ with a linear scheduler decaying to zero, AdamW, batch size 128, and 64 GPU hours. Backbones are MiniCPM5-2B, Ministral-3-8B, and Qwen3.5 at 4B, 9B, 27B, and 122B-A10B. Evaluation covers the SkillGym test set (a held-in subset of unseen tasks with seen skills and a held-out subset of unseen tasks with unseen skills), SkillEval, SkillsBench, and Skill-Use-Bench, using MiniSwe-Agent as the evaluation harness with only a bash tool; skill names and descriptions are placed in the system prompt while detailed contents must be disclosed by the agent itself. Reported metrics are overall task success on SkillGym, mean normalized reward on SkillEval and SkillsBench, and the SU score on Skill-Use-Bench.
Why This Matters
Impact on research. The paper reframes skill-use training as a data-generation problem: instead of mining skills from an agent's own experience in a few environments, it starts from public skills and builds verifiable environments around them. It also contributes a structural analysis of where transfer succeeds and where it fails—training teaches agents to consult skills more than to apply their methods—and shows that a reviewer-based quality gate adds training signal beyond execution checks alone.
Real-world applications:
- Training coding and software-engineering agents that must interpret and apply reusable skill packages rather than improvise.
- Building agents for document processing, finance, science, and marketing workflows, domains the authors list as covered by community-written skills.
- Producing training data for existing agent harnesses such as Claude Code, Codex, and OpenClaw, where skills are already standard.
- Improving skill retrieval and orchestration behavior in deployed agents, since the largest measured change is agents actually reading the skill they are given (28% to 96%).
Industry relevance. The recipe is cheap relative to pretraining: 64 GPU hours of finetuning on 19k trajectories from open-weight teachers. The result that a 9B finetuned model outperforms the 397B Qwen3.5-397B-A17B on two benchmarks is a direct argument for data and environment construction over scaling alone. The released code and data at the project's GitHub repository lower the barrier to extending the pipeline to new skill collections and domains.
Future Directions
- Method- and tool-centric data construction. Gains on SkillsBench concentrate on skills that supply reference documentation, while skills built around a domain method or a bundled tool improve little. The authors point to method and tool-centric tasks as a target for data construction.
- Reinforcement learning on SkillGym environments. The paper trains only with supervised finetuning; its verifiers already provide outcome rewards, making RL a natural next step.
- Making the four-profile task mix pay off. At the token budgets tested, the four-profile mix matched or slightly trailed procedural-only training, so whether profile diversity raises average performance at larger scale remains open.
- Extending transfer and closing the consult-versus-apply gap. Held-out skill transfer works, but the analysis suggests training mostly strengthens consultation rather than application; understanding and closing that gap is an open question.
Target Audience
Researchers and engineers working on LLM agents, agent training data synthesis, and reinforcement or supervised post-training for tool and skill use; benchmark designers interested in verifiable outcome checks; and practitioners deploying skill-based agent harnesses who want to know whether training, rather than prompting alone, improves skill use. Readers without a background in agent harnesses or finetuning will find the pipeline description accessible but the evaluation details dense.
Authors’ abstract
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training