Skip to content
AI.info

Research

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents Overview Research area: Agentic AI post-training — automatic synthesis of executable tasks and their environments fo

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents
arXiv
2609.33772
Published
2026-09-27
Authors
Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu, Canwei Li, Hongjie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Mu Chuan

AI summary

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

Overview

Research area: Agentic AI post-training — automatic synthesis of executable tasks and their environments for training general-purpose tool-using agents.

Technical level: Advanced. The paper assumes familiarity with agent rollouts, reinforcement/post-training pipelines, supervised fine-tuning, and rubric-based evaluation.

Scope (one sentence): The paper proposes Skill2Env, a framework that turns reusable "skills" into complete executable task environments by first deciding which agent capabilities should be challenged, and then uses solver feedback to make those tasks progressively harder.

What This Paper Is About

Training agents to use tools over many steps requires executable environments, but building them by hand does not scale and building them automatically is hard. A skill (a package of domain knowledge, procedures, and tool-use instructions) provides a starting point, but there is a large gap between what a skill contains and a concrete, genuinely challenging task with a working workspace and evaluator. Skill2Env closes that gap by letting the capability demands the training process wants to exercise drive how the task and its environment are constructed, rather than letting the skill alone determine them.

Key Contributions

  1. Skill2Env, a capability-oriented synthesis framework. It organizes agent capability demands along five dimensions (environment understanding, planning, skill usage, long-horizon consistency, error recovery), represents them as reusable difficulty patterns, and instantiates them into task blueprints that jointly specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide joint construction of the task instruction, execution substrate, workspace, and rubric-based evaluator. Applying the full pipeline yields 2,963 executable tasks.

  2. Iterative Task Hardening. A procedure that uses solver execution evidence to detect tasks the solver already handles well, then strengthens existing difficulty-pattern instantiations or adds new compatible patterns, and revises the blueprint and environment accordingly. Generalizable challenges discovered this way can be abstracted into new difficulty patterns and added to the pattern pool.

  3. Empirical demonstration of transfer. Supervised fine-tuning on 1.5K high-scoring trajectories generated from Skill2Env environments produces consistent gains across seven agent benchmarks, improving the unweighted average from 36.6 to 45.0 (+8.4 points).

  4. Skill curation at scale. More than 47K public skill folders were downloaded from ClawHub and screened against three criteria (Practical Utility, Runtime Compatibility, Setup Feasibility), retaining more than 3K skill folders for synthesis.

Main Findings

  • Consistent gains across seven benchmarks with only 1.5K trajectories. Fine-tuning Qwen3.6-35B-A3B raised the unweighted average from 36.6 to 45.0 (+8.4 points). Per-benchmark changes: Terminal-Bench 2.1 +13.5 (44.9 to 58.4), SWE-bench Multilingual +7.7 (63.3 to 71.0), SkillsBench +14.34 in the text (shown as +14.3 in the results table, 32.5 to 46.9), Claw-Eval +5.51 (55.8 to 61.3), τ³-Banking +4.47 (10.7 to 15.1), AutomationBench +7.34 (10.3 to 17.7), VitaBench +5.87 (38.9 to 44.8).

  • The gains are not benchmark-specific. The paper reports no benchmark-specific optimization; environments are organized around general capabilities, and gains appear across seven benchmarks with different domains, interfaces, and task structures.

  • The fine-tuned 35B model beats a much larger in-family model on six of seven benchmarks. Skill2Env exceeds Qwen3.5-397B-A17B on six of the seven benchmarks despite using the smaller Qwen3.6-35B-A3B backbone. Frontier references (GPT-5.4, Claude Opus 4.6, Gemini-3.1 Pro, DeepSeek-V4-Flash-0731, GLM-5.2, Kimi-K2.6, Qwen3.8-27B) remain stronger on several benchmarks, which the authors describe as substantial remaining headroom.

  • SkillsBench improvements span domains and skill-loading conditions. Under OpenHands, gains are 14.3 percentage points with skills on and 10.4 with skills off. With skills on, gains appear in seven of eight domains (finance unchanged); with skills off, all eight domains improve. Median... rather, the leading skills-on gains are media (+46.7), mathematics and operations research (+26.9), and software engineering (+23.4) percentage points.

  • Qualitative behavior changes match the targeted capabilities. Trajectory analysis reports better coordination of dependent processing stages in media production (planning, long-horizon consistency), more complete handling of content distributed across document structures in office tasks (environment understanding), and better adaptation of guidance/helper scripts plus more effective use of execution feedback (skill usage, error recovery).

  • Improvements hold across four Terminal-Bench 2.1 harnesses. Scores rise from 46.1 to 57.3 with pi (+11.2), from 44.9 to 58.4 with Terminus-2 (+13.5), from 38.2 to 55.1 with OpenCode (+16.9), and from 40.5 to 52.8 with Claude Code (+12.3).

  • Iterative Task Hardening makes tasks measurably harder. On 500 paired task identities hardened in both rounds, with DeepSeek-V4-Flash under the pi harness, the full-credit rate (r = 1) falls from 48.40% to 24.80% and 15.40% at iterations 0, 1, and 2, while mean assistant turns rise from 25.91 to 36.54 and 38.85. That is a 33.0-percentage-point reduction in full-credit rate and roughly 50% more interaction steps from iteration 0 to 2.

  • Hardened collections provide better training supervision at a fixed budget. Holding training-set size and quality threshold fixed (500 trajectories per iteration, the same r > 0.9 filter), the seven-benchmark mean rises from 40.58 to 42.51 and 42.87 for iterations 0, 1, and 2 — a 2.29-point improvement from iteration 0 to 2.

  • The FACET comparison is limited. FACET is included as a prior skill-based environment synthesis baseline, but the only reported number is FACET-Terminal-Qwen3.5-27B at 47.6 on Terminal-Bench, with the remaining cells not reported; the paper states FACET results are source-reported.

  • Pattern pool and hardening details. The difficulty pattern pool is initialized with 100 reusable patterns across five capability dimensions. The hardening trigger threshold is set to 0.7, so tasks scoring above it undergo further hardening and the rest are retained. Kimi-K3 is used for environment synthesis and execution diagnosis, Qwen3.6-35B-A3B is the fixed solver during hardening, Qwen3.5-397B-A17B is the LLM rubric judge, and DeepSeek-V4-Flash is the teacher generating trajectories.

Methodology in Plain English

The pipeline has four moving parts:

1. Curate skills. Skills are downloaded and filtered down to ones that are genuinely useful, run in a Linux sandbox, and can actually be set up with their dependencies, data, and services.

2. Decide what the task should train. Rather than asking "what task can this skill do?", the system asks "which agent capabilities do we want to stress?" It uses five capability demands and a library of 100 reusable difficulty patterns — concrete recipes for adding difficulty, such as splitting facts across multiple sources, making later steps depend on earlier outputs, limiting a resource budget, or making a skill procedure valid only under preconditions that must be checked. From these, it selects a compatible subset for the skill.

3. Build against a blueprint. The blueprint is an author-side contract that pins down the objective, the challenges, the facts available, what the solver must discover versus infer, and the acceptance criteria. It drives construction of four artifacts: the task instruction, the execution substrate (tools and runtime), the initial workspace (built from synthesized materials plus retrieved real-world materials), and a rubric-based evaluator. The evaluator assigns weights summing to 1 across rubric items; programmatically checkable items typically return binary scores, while the rest go to an LLM judge that can award partial credit. The task reward is the weighted average of rubric results. A validation agent then checks whether task conditions and rubric results agree with the blueprint and execution evidence; issues that require changing the task contract trigger coordinated blueprint-and-component updates, implementation-only issues are fixed against the existing blueprint, and evidence traced to inconsistent task conditions or evaluator errors is excluded from diagnosis so hardening is grounded in valid demands and results.

4. Harden iteratively. The solver runs the task and produces a reward. Tasks scoring above the 0.7 threshold get hardened. A diagnosis agent compares the trajectory against the blueprint to see how the solver handled each pattern, whether intended challenges were bypassed via shortcuts, and where failures remain. A hardening proposal then strengthens existing patterns or adds new compatible ones, the blueprint is revised, and the task is reconstructed. Newly discovered generalizable challenges can be abstracted into new patterns and added to the pool.

Training. Trajectories generated by DeepSeek-V4-Flash on the resulting 2,963 tasks are filtered with r > 0.9, retaining 1.5K. Qwen3.6-35B-A3B is fine-tuned for 5 epochs with global batch size 32 using Megatron, peak learning rate 1e-5, minimum learning rate 1e-6, cosine decay, and a 10% warmup fraction.

Evaluation protocol. Seven benchmarks: Terminal-Bench 2.1 (Terminus-2, one trial per task, 10,800 s timeout), SWE-bench Multilingual (all 300 instances, mini-SWE-agent, one trial, 7,200 s timeout), SkillsBench (OpenHands with skills enabled, Avg@3 task reward, 10,800 s timeout), Claw-Eval (199 non-multimodal tasks, native agent loop, three trials, Pass^3), τ³-Banking (Pass^1), AutomationBench (pass rate, version 1.0.6), and VitaBench (mean score).

Why This Matters

Impact on research. The paper reframes environment synthesis: instead of maximizing skill or task coverage, it organizes task construction around the capabilities the agent is supposed to exercise. If that framing holds, it offers a way to keep generating environments that stay challenging as agents improve, which is the bottleneck for continued post-training progress. It also connects skill libraries to executable training data, giving a use for the large body of published skills.

Real-world applications:

  • Agent post-training pipelines that need a scalable supply of executable tasks with automated scoring, rather than hand-built benchmarks.
  • Document and office workflows, where the reported Docling case study and office-task behavior improvements point to agents that reconcile conflicting evidence and avoid dropping content spread across document structures.
  • Media production pipelines, where the reported largest skills-on gain (+46.7 points) reflects coordination of dependent processing stages while preserving final-output requirements.
  • Operations and engineering toolchains, where capability-focused supervision may transfer to repository-level issue resolution and cross-application automation, matching the reported gains on SWE-bench Multilingual and AutomationBench.

Industry relevance. The headline commercial result is efficiency: 1.5K supervised trajectories and a 35B-parameter model produce gains across seven benchmarks, and the tuned model outperforms a much larger model from the same family on six of seven. For teams with limited compute and no access to frontier closed models, this suggests environment design quality can substitute for scale. The cross-harness consistency (terminal-task gains under pi, Terminus-2, OpenCode, and Claude Code) also matters for deployment, since agent products differ in their execution scaffolding.

Future Directions

  • Whether hardening can continue indefinitely. The paper hardens for two iterations; how many rounds remain useful before tasks become unsolvable, degenerate, or saturated is not established.
  • How far the transfer extends. Gains are shown on seven benchmarks plus four terminal harnesses; generalization to entirely different modalities, interfaces, or long-running real deployments is not reported.
  • Cost and throughput of the pipeline. The paper does not report wall-clock time, compute cost, or token usage for synthesis, diagnosis, or hardening, so the practical economics relative to manual environment authoring are not quantified.
  • Compositional and multi-agent environments. The paper notes related work extends synthesis through skill teams and graphs, and asks how capability demands scale beyond single-skill, single-solver settings; this is not addressed.
  • Sensitivity to the hardening threshold and pattern pool. Only a 0.7 trigger threshold and a 100-pattern initial pool are reported; how results change with different values is not reported. The paper's appendix tables describing difficulty patterns are truncated in the available content, and Appendices C, D, and E (diversity analysis, per-benchmark hardening scores, and the Docling case study) are referenced but their contents are not included here.

Target Audience

Researchers and engineers working on agent post-training, synthetic data generation, and executable environment construction, especially those who build training data from reusable skill or tool libraries. It is most useful to readers who already understand agent rollouts, SFT pipelines, and rubric-based evaluation, and who want a concrete recipe for making synthesized tasks harder on purpose rather than by random variation. Benchmark designers and evaluation researchers will also find the cross-harness and domain-level breakdowns useful, though readers seeking implementation cost accounting or full appendix detail will need the complete paper, since portions of the appendices are not present in the available content.

Authors’ abstract

Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.

Read the original paper