Research
CompoWorld: Compositional Environment Scaling for General Agents
Overview Research area: LLM agents, automated environment/task synthesis, tool-use training (SFT and reinforcement learning), compositional generalization. Technical level: Advanced — the paper uses P

- arXiv
- 2609.33665
- Published
- 2026-09-27
- Authors
- Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan
AI summary
Overview
- Research area: LLM agents, automated environment/task synthesis, tool-use training (SFT and reinforcement learning), compositional generalization.
- Technical level: Advanced — the paper uses POMDP/product-environment formalism, reward derivations, and GRPO policy optimization math.
- Scope: The paper introduces CompoWorld, a framework that builds 448 reusable executable services exposing 10,130 tools and composes them into verified cross-service tasks to train a general agent (Qwen3.6-35B-A3B) with 3K SFT trajectories and 1K RL tasks.
What This Paper Is About
Existing automated environment generators mostly create tasks inside a single environment, but real workflows require an agent to carry information and actions across several systems. CompoWorld addresses this by treating a finite library of reusable, independently executable services as building blocks, then composing them through task-specific dependency graphs so that the agent must move information from one service to another to succeed. The goal is to make the composition of services — not just the count of environments — an explicit, controllable dimension of training-data scaling, and to use those verified composed tasks for both supervised fine-tuning and reinforcement learning.
Key Contributions
- Compositional environment scaling. A framework that composes independently executable, stateful services into product environments with namespaced tools, so the same implementations support many different application workflows. The framework is built from 448 services and 10,130 tools.
- Verified service construction with selective world-model simulation. Coding agents inside a sandbox turn public Model Context Protocol (MCP) specifications into typed Python services with Pydantic-validated states, generate their own test suites, and are then checked by a separate adversarial validation pass; tools that cannot be reliably implemented are handled by an LLM world model instead of being discarded.
- Automated cross-service task generation and verification. A random-walk procedure samples services and connects them into a service-level dependency graph, with walk length L as the main difficulty control, and the resulting tasks are verified by executing a reference solution.
- Completion-Focused Rubric Reward for RL. A rubric-based reward that weights each criterion by its group pass rate, w_j = λ + (1 − p_j), so criteria that are rarely satisfied receive more learning signal, combined with GRPO.
Main Findings
- Average gain over the backbone: CompoWorld improves over Qwen3.6-35B-A3B by 9.17 points on average across eight agent benchmarks.
- AutomationBench: task success rises from 10.33% to 32.33% (+22.00 points), exceeding GPT-5.4 (27.67%), Gemini-3.1 Pro (28.17%), and GLM-5.2 (26.33%), and approaching DeepSeek-V4-Flash (36.33%). It also surpasses Claude Opus 4.6 (25.50%).
- Leading the 35B-A3B agent-specialized group: CompoWorld leads all six compared agent-specialized 35B-A3B models on AutomationBench, beating the strongest reported baseline, Occamy-1.0 (27.60), by 4.73 points.
- Other benchmark gains: SkillsBench +15.19 (32.52 to 47.71), VitaBench +10.50 (38.94 to 49.44), DeepPlanning +9.17 (26.04 to 35.21), τ³-Banking +5.84 (10.65 to 16.49), ALE +5.82 (7.77 to 13.59), WildClawBench +3.18 (44.28 to 47.46), and VitaBench 2.0 +1.62 (34.47 to 36.09).
- Weaker transfer area: The small VitaBench 2.0 gain (+1.62) suggests more limited transfer to personalized, long-term assistance; frontier models still lead on τ³-Banking, DeepPlanning, and VitaBench 2.0.
- Domain-level partial credit on AutomationBench 1.0.6: the mean score rises from 41.94 to 72.68. Gains are +49.40 (HR), +32.51 (Marketing), +31.45 (Sales), +26.56 (Finance), +23.00 (Support), and +21.49 (Operations). CompoWorld exceeds both GPT-5.4 and GLM-5.2 in five of six domains, with Support as the exception, and Sales remains the lowest-scoring domain at 58.93.
- Composition space size: with n = 448 services, composing 5 yields approximately 1.47 × 10^11 possible sets and composing 7 yields approximately 6.86 × 10^14.
- Composed beats single-environment training: at a fixed budget of 1K trajectories, composed-environment SFT raises τ³-Banking from 9.62 to 17.53 (+7.91) and AutomationBench from 12.83 to 33.33 (+20.50) versus single-environment SFT; single-environment training even degrades slightly relative to the base model on τ³-Banking and DeepPlanning.
- SFT data scaling: all four tested benchmarks improve over the backbone with as few as 100 SFT examples. At 3K, τ³-Banking rises from 10.65 to 17.87 and DeepPlanning from 26.04 to 35.02. AutomationBench reaches 33.33% at 1K and 32.67% at 3K; SkillsBench reaches 44.23 at 500, peaks at 45.13 at 1K, and declines to 40.91 at 3K, still above the backbone's 32.52.
- RL dynamics: the smoothed mean trajectory reward reaches about 0.69 within the first ten steps, settles around 0.65–0.67, reaches about 0.75 at step 80, and ends near 0.79 at step 150. RL improves on SFT by 1.2 points on average across the eight benchmarks but declines slightly on τ³-Banking; the paper states the overall improvements are primarily driven by SFT.
- Domain breadth of the service corpus: software development and DevOps is the largest domain at 22%, followed by productivity and collaboration at 12%; no single domain exceeds a quarter of services, and thirteen domains each contribute a notable share.
Methodology in Plain English
The pipeline has three stages.
Building services. The team crawled machine-readable MCP specifications, kept those for relatively self-contained applications, and normalized them into a unified function-calling format. A coding agent working inside an isolated sandbox then inferred each service's entities, wrote a typed Pydantic state model, implemented every tool as an operation over that state, and generated a test suite — iterating until both successful-use and error-handling cases passed. Because self-generated tests can share the implementation's blind spots, a separate agent session derived adversarial test cases from the tool specifications, including boundary conditions. Tools failing these checks were quarantined and not used as verified deterministic transitions. For long-tail tools that depend on external systems or exceed the coding agent's ability, an LLM world model predicts both the state update and the observation, and those predicted updates are instantiated through the same Pydantic models so they stay type-valid. Each service ships as its own environment with persistent state and namespaced tools, and all services share one execution convention and unified JSON response format, which is what makes them directly composable.
Generating tasks. A task is a tuple of instruction, initial joint state, service-level dependency graph, a reachable reference goal, and a verifier. The system samples K services from the pool, then uses a random walk that must cover every selected service, never visit the same service twice in a row, and revisit at least one service — so each edge encodes an information dependency and revisits force the agent to re-read a value that has gone stale. Walk length L (sampled from 7 to 12 in the experiments) is the main difficulty dial; K and the sampled domain are secondary controls. A coding agent then runs four phases: explore the services, design a concrete scenario with a validated initial state and a structured rubric, probe by solving the task so the resulting state defines the goal and each rubric criterion becomes an executable final-state check, and submit the final artifacts. Verification requires that the reference solution is reachable via real tool calls and that the verifier accepts it, while still allowing other valid trajectories that differ in unrelated state fields.
Training. The policy is first fine-tuned on verified successful trajectories (3K for three epochs, global batch size 32, learning rate 10^-4), then trained with RL on 1K tasks using GRPO (Adam, learning rate 10^-5, clipping parameters 0.20 and 0.28, N = 8, λ = 0.01, zero KL and entropy coefficients). GRPO normalizes rewards within each rollout group to compute advantages, assigns that outcome advantage to all agent-generated tokens, and excludes tool observations from the loss. The reward itself is the Completion-Focused Rubric Reward: each criterion is weighted more heavily when few rollouts in the group satisfy it, so the signal concentrates on unmet requirements while still giving partial credit. Environment and task synthesis used GLM-5.3, trajectories were generated with DeepSeek-V4-Flash (0731), and Qwen3.5-397B-A17B served as the LLM judge where a judge was needed. The agent sees four meta-tools: list_services, list_tools, describe_tool, and call_tool.
Why This Matters
Impact on research. The paper reframes environment scaling as a compositional problem: rather than implementing a new environment for every training example, a fixed library of reusable services is recombined, so adding one SFT example can also add one training environment. It connects environment synthesis to compositional generalization and provides a reward-shaping mechanism aimed specifically at the gap between high partial rubric scores and low full-completion rates.
Real-world applications (drawn from the paper's benchmarks and examples):
- Cross-application business workflows — the AutomationBench domains evaluated here are Finance, HR, Marketing, Operations, Sales, and Support, with HR showing the largest gain (+49.40 points).
- Banking customer support requiring knowledge retrieval and policy-compliant tool use, as measured by τ³-Banking.
- Multi-system information handoffs such as the paper's example of querying a Snowflake database for overdue tickets, consulting a PDF operating manual, and then emailing managers and customers.
- Personalized, proactive assistance over long-term interactions, as tested by VitaBench 2.0, and long-horizon professional tasks with verifiable outcomes, as tested by Agents' Last Exam (ALE).
- Reusable-skill agents, corresponding to the SkillsBench evaluation where the largest gain (+15.19) was observed.
Industry relevance. The corpus is deliberately broad rather than concentrated — no domain exceeds a quarter of services — and is built from public MCP specifications, which are the interface format many tool ecosystems already use. That makes the approach relevant to teams wanting training data for agents that coordinate email, chat, calendars, databases, and business software, and the 35B-A3B comparisons give a concrete reference point for small-scale specialized agent models. The paper also reports where the approach falls short: a 32.33% AutomationBench pass rate means completing every requirement stays difficult even when partial progress is substantial.
Future Directions
- Growing the service library as a scaling axis. The paper notes that adding a service to a library of n services creates binom(n, K−1) additional candidate compositions of size K, so expanding the library complements generating more compositions from the existing pool. The stated practical challenge is selecting compatible services, constructing substantive cross-service dependencies, and verifying the resulting tasks.
- Improving verification and coverage for long-tail tools. Tools that fail adversarial checks are quarantined, and tools that cannot be faithfully implemented fall back to world-model simulation; the paper discusses environment quality and the limited scope of simulation in its appendix, leaving open how far supervised trust in simulation can extend.
- Closing the full-completion gap. The 32.33% AutomationBench pass rate and the observation that "completing every requirement remains difficult" point to harder rewards or task curricula; Sales remains the lowest-scoring domain at 58.93, which the paper attributes to room for strengthening the dependencies and constraints in those workflows.
- Making RL more reliably beneficial and broadening transfer. RL improves on SFT by 1.2 points on average but declines slightly on τ³-Banking, the gain on VitaBench 2.0 is only +1.62, and improvements are primarily driven by SFT — so understanding when RL helps and how to transfer to personalized long-term assistance remains an open question.
Target Audience
Researchers and engineers working on LLM agents, tool use, and training-environment synthesis; practitioners building agents that span multiple enterprise applications (email, chat, calendars, databases, CRM); and readers interested in compositional generalization or reward design for agentic reinforcement learning. The paper assumes familiarity with POMDPs, GRPO, and tool-calling interfaces, so readers without a machine learning research background will find the formal sections in Sections 3 and 4.3 demanding.
Authors’ abstract
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.