Skip to content
AI.info

Research

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

Overview Research area: Tool-use post-training for large language model (LLM) agents — specifically, how to build the surrounding infrastructure that turns raw model capability into reliable tool-usin

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
arXiv
2609.36887
Published
2026-09-29
Authors
Bo Mao, Hang He, Linting Wang, Lizhi Lin, Maosen Zhou, Guanming Liu, Jinxiu Liu, Tianyu Huai, Chaoyun Zhang, Bingxuan Li, Kepeng Lei, Guanting Dong, Zhou Shao, Rui Zheng, Hang Yan, Jie Zhou, Chengcheng Wan, Tao Gui, Liang He, Xipeng Qiu

AI summary

Overview

Research area: Tool-use post-training for large language model (LLM) agents — specifically, how to build the surrounding infrastructure that turns raw model capability into reliable tool-using behavior.

Technical level: Advanced. The paper assumes familiarity with LLM post-training, reinforcement learning-style credit assignment, agent harnesses, and agent evaluation benchmarks.

Scope: The paper introduces WEFT, a whole-system framework for constructing, evolving, and stably post-training general-purpose tool-using agents, and reports benchmark comparisons against environment-scaling baselines.

What This Paper Is About

Recent work on scaling tool-use post-training has focused almost entirely on synthesizing executable environments, but an environment is only one piece of a larger agentic interaction system that also includes the task, the agent harness, and the evaluator. The authors argue that scaling environments in isolation does not reliably improve model performance, because useful learning signals depend on all four components interacting coherently. WEFT is their attempt to scale and improve the whole system together rather than one component at a time.

Key Contributions

  1. A whole-system framing of tool-use post-training. The paper reframes the problem from "generate more environments" to "build and evolve a coherent system of environment, task, agent harness, and evaluator," arguing that this coherence is what produces dependable learning signals.

  2. Scalable agentic interaction system construction. WEFT scales construction across three dimensions: environment breadth, task complexity, and interaction diversity — rather than broadening only the environment axis.

  3. Execution-driven self-evolution. The framework iteratively uses execution traces and state evidence to attribute failures to the specific component responsible, revises that component, then runs fresh rollouts to evaluate the change and generate evidence for the next evolution round.

  4. Mechanisms for stable post-training at scale. Prefix-preserving sampling retains verified progress, atomic-turn credit assignment localizes learning signals, and MegaMCP maintains isolated, recoverable state across concurrent rollouts that share tool services.

Main Findings

  • WEFT outperforms environment-scaling baselines at matched model size. WEFT-8B and WEFT-14B beat all evaluated matched-size environment-scaling baselines on BFCL V4, τ²-Bench, and Claw-Eval.

  • The 14B model shows substantial gains over a comparable baseline. WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points across those three benchmarks respectively. The abstract does not report absolute scores for either model.

  • Gains extend to long-horizon workflows. WEFT-35B-A3B carries the improvements into more challenging long-horizon workflow benchmarks, specifically Toolathlon-Verified and AutomationBench. The abstract states that the gains extend but does not quantify them for this model.

  • Effectiveness holds across multiple models and benchmarks. The authors describe extensive experiments spanning various models and benchmarks, though the abstract gives no dataset sizes, training compute figures, or ablation results.

Methodology in Plain English

The authors start from the observation that an agent's performance depends on four things working together: the environment it acts in, the task it is given, the harness that connects the model to tools, and the evaluator that judges success. If you improve only the environment, the other three can still be the bottleneck.

WEFT therefore scales all four at once, along with the variety of interactions the agent encounters. Once the system is running, it learns from its own executions: when something fails, the framework looks at the execution trace and state evidence to figure out which component is at fault, then revises that component. New rollouts test whether the revision helped, and their evidence feeds the next round — a self-improvement loop aimed at the system rather than only the model.

Two sets of engineering problems arise when training at this scale, and WEFT handles both. On the optimization side, prefix-preserving sampling keeps verified progress intact rather than discarding it, and atomic-turn credit assignment narrows down which action deserves credit or blame. On the execution side, MegaMCP gives each concurrent rollout its own isolated, recoverable state even though many rollouts are hitting the same shared tool services. The abstract presents these as reliability mechanisms, but does not detail their internal design.

Why This Matters

Impact on research. The paper challenges a common assumption in the agent-training literature — that more environments automatically mean better agents — and proposes that the bottleneck is systemic coherence. If the argument holds, it shifts effort toward jointly designing tasks, harnesses, and evaluators alongside environments, and toward self-evolution loops that can diagnose and repair the weakest component.

Real-world applications:

  • Software and IT automation agents that must call APIs, run code, and recover from errors across long multi-step workflows.
  • Enterprise workflow agents handling processes where partial progress must not be lost and where concurrent execution against shared services is the norm.
  • Customer-facing assistants that need reliable tool calls (order lookups, bookings, account changes) rather than plausible-sounding but unexecuted actions.
  • Research and data pipelines where agents orchestrate many tools over extended horizons and failures need to be traced back to a specific stage.

Industry relevance. The claimed gains at 8B and 14B parameter scales matter commercially: smaller models are cheaper to serve, and a training method that closes the gap with larger or better-scaffolded baselines changes the cost/performance tradeoff for deployed agents. The infrastructure concerns WEFT addresses — state isolation across concurrent rollouts, credit assignment over long trajectories, preserving verified progress — are exactly the engineering obstacles teams hit when they try to move agent training from demos to production.

Future Directions

  • Verify the whole-system claim against component scaling. The abstract argues environments alone are insufficient, but it does not report an ablation isolating how much each of environment, task, harness, and evaluator contributes to the gains.
  • Test whether self-evolution converges or drifts. Iterative failure attribution and component revision could accumulate inconsistent changes; the abstract does not say how many evolution rounds were run or how stability was measured.
  • Extend evaluation to longer and more open-ended horizons. Gains are claimed on Toolathlon-Verified and AutomationBench for the 35B-A3B model, leaving open how the approach scales to tasks far beyond current benchmark horizons.
  • Generalize beyond the tested model families and benchmarks. Whether the same system-construction and reliability mechanisms transfer to other agent domains, tool ecosystems, or evaluation regimes is not established in the abstract.

Target Audience

Researchers and engineers working on LLM agents — particularly those building tool-use post-training pipelines, agent harnesses, or synthetic environment generation. It is also relevant to ML infrastructure engineers concerned with concurrent rollout execution and state management, and to practitioners deciding how much of their agent-quality problem can be solved by more data versus better system design. Readers without background in post-training and agent evaluation will find the advanced sections difficult, though the core argument about whole-system coherence is accessible to a general technical audience.

Authors’ abstract

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, $τ^2$-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.

Read the original paper