Research
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
Overview Research area: Artificial Intelligence — LLM agents, agentic reinforcement learning, and skill/memory library management. Technical level: Advanced (assumes familiarity with reinforcement lea

- arXiv
- 2610.09832
- Published
- 2026-10-07
- Authors
- Yuyao Ge, Yiwei Wang, Yuchen He, Baolong Bi, Lingrui Mei, Jiayu Yao, Lizhe Chen, Shenghua Liu
AI summary
Overview
Research area: Artificial Intelligence — LLM agents, agentic reinforcement learning, and skill/memory library management.
Technical level: Advanced (assumes familiarity with reinforcement learning, GRPO-style policy optimization, and retrieval-augmented LLM agents).
Scope: The paper proposes SkillForge, a method that treats an agent's skill library as a dynamic population managed by a fitness-driven lifecycle (trial, active, stable, retired) with LLM-guided mutation, so that skills and the policy co-evolve during online RL, plus the associated SkillFurnace dataset.
What This Paper Is About
Skill libraries give LLM agents reusable instructions, examples, and applicability conditions that are retrieved into the context on demand, and pairing them with reinforcement learning is a natural way to focus exploration while shaping the policy. The problem is that existing skill-augmented RL methods treat the library as append-only: once a skill is added it persists, so low-quality, outdated, or now-harmful skills keep polluting the agent's context even as the policy improves. SkillForge addresses this by actively forging the library — pre-filtering weak skills before training, then retiring, stabilizing, demoting, and mutating skills as the policy evolves, so the active set stays compact and aligned with the agent's current capability.
Key Contributions
- SkillForge and the skill lifecycle. A fitness-driven lifecycle that moves every skill through
trial,active,stable, andretiredstates, augmented with LLM-guided mutation, replacing the append-only accumulation used by prior skill-augmented RL. Skills and model co-evolve across training iterations. - SkillFurnace dataset. A three-part bundle of 5k+ annotated records (5,852 total records across the three environments) spanning retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories.
- Empirical validation. Across three interactive agent benchmarks spanning embodied control, web navigation, and Search-Augmented QA, SkillForge surpasses the strongest baseline on every environment while keeping the library compact, delivering up to 7.8% relative improvement in aggregate success rate.
- Formalization of delayed obsolescence. The paper names and formalizes the failure mode in which a skill that once reached the
stablestate later falls below the retirement threshold as the policy matures.
Main Findings
- Highest aggregate success rate. SkillForge achieves 92.4% on ALFWorld and 78.4% on WebShop, improving over SkillRL by 2.8% relative on ALFWorld and 7.8% relative on WebShop. It outperforms the closed-source reference points GPT-4o (48.0% ALFWorld All; 31.8 WebShop Score, 23.7 Succ.) and Gemini-2.5-Pro (60.3% ALFWorld All; 42.5 WebShop Score, 35.9 Succ.) by a wide margin.
- Strongest per-subtask gain. The largest per-subtask relative improvement over SkillRL is on ALFWorld's Look at 21.3%; the Clean subtask remains within 0.2% of SkillRL, within sampling noise.
- Search-Augmented QA. SkillForge reaches 48.7 ± 0.3% micro-averaged accuracy on the full test set, the highest overall among all methods.
- Faster, higher convergence. SkillForge and SkillRL both climb from ~60% to ~70% between steps 50 and 65, but SkillForge separates beyond step 75 as the forging lifecycle takes effect, reaching ~90% while SkillRL oscillates in the ~80–88% range for the rest of the training budget.
- Every lifecycle stage matters. Ablations show the highest success rates come from the full configuration: 92.4 / 78.4 / 48.7 (ALFWorld / WebShop / Search-Aug. QA). Removing the entire lifecycle drops to 89.8 (▼2.6), 72.6 (▼5.8), and 46.3 (▼2.4); removing pre-retirement drops to 90.7 (▼1.7), 74.9 (▼3.5), 47.5 (▼1.2); removing retirement drops to 91.4 (▼1.0), 75.9 (▼2.5), 47.9 (▼0.8); removing mutation drops to 91.8 (▼0.6), 76.8 (▼1.6), 48.2 (▼0.5). The paper's text describes the largest drop as peaking at 7.4% on WebShop and lists pre-retirement per-stage relative drops of 1.8%, 4.5%, and 2.5%, which differ from the values printed in Table 3; the table values are given above.
- Bigger libraries are not better. Without retirement the library grows to over 130 skills on ALFWorld (132 in the table) yet SkillForge outperforms it with a compact set of 100. On WebShop, enabling pre-retirement and mutation lifts the static 72.6% baseline to 75.9%; adding retirement trims the library from 121 skills to 95 and closes the remaining 3.2% relative gap to the full 78.4%.
- Sub-additive components. The single-component drops consistently sum to more than the joint w/o forging drop, consistent with sub-additive interactions among the three components.
- Delayed obsolescence is empirically visible. In SkillFurnace, retired skills peak at an average fitness of 0.67, close to the stabilization threshold δ_stable of 0.7, before dropping by 0.28 at retirement, and 57.2% first reach the
stablestate before this decay. - Library dynamics reach equilibrium. On ALFWorld, 200 training steps grow the library to 100 non-retired skills, with 132 created and 32 retired; the cap binds near step 170, after which mutations and retirements balance into dynamic equilibrium. At the final checkpoint 76 skills are
stableand 24 areactiveortrial. - Two opposing flows, not monotone improvement. On ALFWorld the sub-δ_retire fraction shrinks from 7.5% to 2.0% across training, while mutation expands the middle band [δ_retire, δ_stable) from 18.9% to 22.0% as children enter with a neutral prior at 0.5; the stable regime retains most of the library and mean fitness stays near 0.73.
- Dominant failure mode. Procedural rigidity emerges as the dominant mode by which once-useful skills turn into active interference.
Methodology in Plain English
SkillForge runs in three stages.
First, a pre-RL evaluation phase tests the initial seed skill library using the base model's own rollouts. Each skill's proto-fitness is its success count divided by its usage count, and skills below a conservative threshold δ_pre that have been used at least N_min times are pre-retired. Successful trajectories that never relied on a subsequently retired skill form the SFT dataset, and fine-tuning the base model on that data produces the policy that initializes RL and also serves as the KL reference policy.
Second, skill fitness tracking continues the counters from the pre-retirement phase rather than resetting them, so proto-fitness is preserved. Fitness is the empirical success rate once a skill has at least N_warm usages, and a neutral default prior otherwise. Credit is assigned at the episode level: every retrieved skill in a successful episode gains a success, every retrieved skill in any episode gains a usage.
Third, the forging cycle runs every 10 training steps and applies a lifecycle rule. trial skills promote to active after N_promote usages; active skills stabilize to stable when fitness reaches δ_stable with at least N_stable usages; stable skills whose fitness falls below δ_demote are demoted back to active; and active skills below δ_retire with enough usages are retired. Only active skills can be retired, and distilled (generation-0) skills face a higher usage bar (N_protect). This hysteresis prevents premature removal of new skills and protects proven skills from transient dips. Borderline active skills with fitness in [δ_low, δ_high] and at least N_mutate usages are mutation candidates; up to K_mutate parents are sampled with probability proportional to 1 − fitness, and a teacher model from a different family (Kimi-K2.5) rewrites them using their failure trajectories. Each child enters as trial with generation incremented. A total skill cap S_max bounds cost.
Training uses GRPO on top of the verl framework with Qwen2.5-7B-Instruct as the base model, 200 steps, learning rate 1 × 10⁻⁶, clipping range ϵ = 0.2, KL penalty coefficient β = 0.001, rollout temperature 1.0, and group size G = 8, with K_mutate = 5 mutations and K_retire = 3 retirements per cycle.
Why This Matters
Impact on research. The paper reframes the skill library from a static asset into a population under selection pressure that co-evolves with the policy, and it names a concrete, measurable failure mode — delayed obsolescence — that monotone memory approaches cannot address. Because fitness is inherently non-stationary under online training, this connects agentic RL to quality-diversity and evolutionary optimization ideas that previously operated offline against a fixed held-out metric. The released SkillFurnace dataset, with human-annotated retirement events and failure categories, gives the community a benchmark for studying skill quality and lifecycle management rather than only end-task success.
Real-world applications.
- Personal and enterprise agents that must absorb personalized workflows and post-deployment tooling conventions that fall outside a model's fixed knowledge cutoff.
- Coding and software-engineering assistants, where Claude Code and OpenClaw are cited as examples of systems that derive capability from modular, reusable abstractions provided at inference time.
- Web-navigation and shopping agents that must follow evolving site conventions.
- Search-augmented question answering systems that retrieve both documents and reusable reasoning skills.
Industry relevance. Any deployment that accumulates procedural memory over time faces the append-only problem: retrieval quality degrades, outdated entries crowd out good ones, and context budgets inflate. A lifecycle with retirement and mutation keeps the active library compact (100, 95, and 85 skills in the reported configurations) while improving task success, which matters directly for inference cost and reliability. The finding that quality control rather than scale drives gains argues against "just add more memory" as a scaling strategy.
Future Directions
- Tuning the lifecycle thresholds and budgets. δ_retire, δ_demote, δ_stable, the mutation band [δ_low, δ_high], N_min, N_warm, N_promote, N_stable, N_mutate, N_protect, K_retire, K_mutate, and S_max are all fixed hyperparameters; the paper does not report sensitivity analyses, and the ablation shows the components interact sub-additively, so joint tuning is an open question.
- Teacher-model dependence. Mutation relies on a teacher from a different model family, motivated by a report that self-generated skills provide negligible or negative benefit. Whether cheaper or self-hosted teachers can substitute is not reported.
- Retirement attribution and failure taxonomy. The dataset includes retirement events with human-annotated failure categories and identifies procedural rigidity as the dominant mode, but the paper does not report whether specific failure categories are more or less recoverable through mutation.
- Generalization beyond the three benchmark families. Results cover ALFWorld, WebShop, and Search-Augmented QA; transfer of the lifecycle to other agentic domains, and the behavior of the method at much larger libraries or longer training budgets, is not reported. Limitations of the approach are not stated in the provided content.
Target Audience
Researchers and engineers working on LLM agents, agentic reinforcement learning, and memory or skill augmentation for language models, particularly those interested in experience accumulation, retrieval-augmented policies, and LLM-guided evolutionary optimization. It is also relevant to practitioners building production agents that accumulate procedural memory over long deployments and need to control library growth and staleness. Readers without a background in policy-gradient RL and group-relative advantage estimation will need to consult the preliminaries in Section 3 first.
Authors’ abstract
Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.