Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
Figures & tables
Figure 1: Performance of CompoWorld and five foundation models on four challenging agent benchmarks.
Figure 2: CompoWorld overview. (a) Verified tools with selective world-model simulation; (b) cross-service task generation and verification; (c) completion-focused rubric rewards for agent training.
Figure 3: Domain counts of 448 environments and 10,130 tools, with shared row colors.
Model
τ3 - Banking
Deep Planning
Vita Bench
Vita Bench2.0
Automation Bench
WildClaw Bench
Skills Bench
ALE
Frontier Closed-Source Models
GPT-5.4
28.52
53.96
47.09
45.99
27.67
58.00
51.70
22.50
Claude Opus 4.6
20.27
55.21
38.31
37.71
25.50
54.60
50.20
13.30
Gemini-3.1 Pro
23.71
47.08
52.69
49.53
28.17
38.70
60.80
17.10
Open-Weight Models
DeepSeek-V4-Flash
30.34
54.79
56.31
41.12
36.33
47.67
53.75
17.48
Table 1: Main results on eight challenging agent benchmarks. Green values show score-point gains over Qwen3.6-35B-A3B.
Model
Finance
HR
Marketing
Operations
Sales
Support
GPT-5.4
67.39
62.05
71.50
77.21
53.56
78.05
Gemini-3.1 Pro
75.26
72.07
78.37
72.26
61.37
74.05
DeepSeek-V4-Flash
60.98
59.69
82.94
82.89
62.44
78.12
GLM-5.2
67.46
65.95
71.70
70.26
54.83
77.19
Kimi-K2.6
49.79
48.85
50.41
64.29
42.27
62.62
Qwen3.5-397B-A17B
33.35
14.36
29.07
35.92
22.97
36.35
Table 2: Domain-level results on AutomationBench 1.0.6 (average score, %). Best scores are bold ; second-best scores are underlined . Green values show score-point gains over Qwen3.6-35B-A3B.
Figure 4: Effect of SFT training-set size on Qwen3.6-35B-A3B across four benchmarks. The zero-sample setting denotes the initial backbone.
Figure 5: Environment composition and RL training. (a) Comparison between single-environment scaling and composed-environment scaling; (b) Mean trajectory reward during the RL stage.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Smoothed evaluation-set scores with and without reward reweighting.
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling based on language models can further push the boundaries of general agents. (i) We first focus on building foundation models for agentic environment simulation. We introduce Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, the first language world models capable of simulating agentic environments covering 7 domains via long chain-of-thought reasoning. Leveraging more than 10M environment interaction trajectories of 7 domains in real-world environments, we develop Qwen-AgentWorld through a three-stage training pipeline: CPT injects general-purpose world modeling capabilities from the state transition dynamics and augmented professional corpora, SFT activates next-state-prediction reasoning, and RL sharpens simulation fidelity through a tailored framework with hybrid rubric-and-rule rewards. To evaluate language world models, we present AgentWorldBench, a comprehensive benchmark constructed from real-world interactions of 5 frontier models on 9 established benchmarks. Empirical results demonstrate that Qwen-AgentWorld significantly outperforms existing frontier models. (ii) Beyond foundation models, we further investigate two complementary paradigms through which world modeling enhances general agents. First, as a decoupled environment simulator, Qwen-AgentWorld supports scalable and controllable simulation of thousands of real-world environments for agentic RL, yielding gains that surpass real-environment training alone. Second, as a unified agent foundation model, world-model training acts as a highly effective warm-up that improves downstream performance across 7 agentic benchmarks. Code: https://github.com/QwenLM/Qwen-AgentWorld
Equipping LLMs with tool-use capabilities via Agentic Reinforcement Learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robust execution environments and the scarcity of realistic training data that captures implicit human reasoning. Existing approaches depend on costly real-world APIs, hallucination-prone LLM simulators, or synthetic environments that are often single-turn or depend on pre-collected documents. Moreover, synthetic trajectories are frequently over-specified, resembling instruction sequences rather than natural human intents, reducing their effectiveness for RL training. We introduce EnvFactory, a fully automated framework that addresses both challenges. EnvFactory autonomously explores and verifies stateful, executable tool environments from authentic resources, and synthesizes natural multi-turn trajectories through topology-aware sampling and calibrated refinement, producing grounded queries with implicit intents. Using only 85 verified environments across 7 domains, EnvFactory generates 2,575 SFT and RL trajectories. Despite using significantly fewer environments than prior work, which are often 5 times more, EnvFactory achieves superior training efficiency and downstream performance, improving Qwen3-series models by up to +15% on BFCLv3, +8.6% on MCP-Atlas, and +6% on conversational benchmarks including τ2-Bench and VitaBench. By fully automating both environment construction and trajectory synthesis, EnvFactory provides a scalable, extensible, and robust foundation for Agentic RL.
Minrui Xu, Zilin Wang, Mengyi DENG +12
LARK, HKUST (GZ) · University of Cambridge · UCL +2
Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model Context Protocol (MCP) and broader agent skills offer a unified interface for connecting agents with scalable real-world services, but training robust agents remains limited by the lack of realistic environments and principled mechanisms for life-long learning. In this paper, we present \textbf{Agent-World}, a self-evolving training arena for advancing general agent intelligence through scalable environments. Agent-World has two main components: (1) Agentic Environment-Task Discovery, which autonomously explores topic-aligned databases and executable tool ecosystems from thousands of real-world environment themes and synthesizes verifiable tasks with controllable difficulty; and (2) Continuous Self-Evolving Agent Training, which combines multi-environment reinforcement learning with a self-evolving agent arena that automatically identifies capability gaps through dynamic task synthesis and drives targeted learning, enabling the co-evolution of agent policies and environments. Across 23 challenging agent benchmarks, Agent-World-8B and 14B consistently outperforms strong proprietary models and environment scaling baselines. Further analyses reveal scaling trends in relation to environment diversity and self-evolution rounds, offering insights for building general agent intelligence.