Scaling Verifiable Environments for Long-horizon Work Agents
Abstract
Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.
Figures & tables
| Model | WildClaw | ClawEval | GDPval | APEX | ALE | ||||
|---|---|---|---|---|---|---|---|---|---|
| Score | Pass@1 | Score | Passˆ3 | Score | Score | Pass@1 | Score | Pass@1 | |
| Frontier Proprietary Models | |||||||||
| Claude Opus 5.0 | 64.5 | 32.4 | 82.3 | 74.2 | 91.0 | 58.5 | 42.7 | 48.3 | 26.7 |
| Claude Opus 4.8 | 64.0 | 30.5 | 81.5 | 72.1 | 89.1 | 46.2 | 34.0 | 46.2 | 24.8 |
| GPT-5.6 Sol | 53.5 | 29.5 | 77.3 | 68.9 | 90.1 | 51.8 | 37.5 | 48.3 | 26.7 |
| Gemini 3.7 Flash | 52.6 | 18.1 | 69.2 | 60.0 | 86.1 | 51.1 | 36.9 | 46.9 | 25.7 |
| ALE | APEX | ||
|---|---|---|---|
| Score | Pass@1 | Score | |
| LLM only | 15.62 | 4.8 | 11.18 |
| Programmatic | 14.89 | 4.8 | 10.92 |
| Hybrid | 22.11 | 11.4 | 15.39 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| # | Domain | Count | % | # | Domain | Count | % |
|---|---|---|---|---|---|---|---|
| 1 | Office Productivity | 1,394 | 8.36 | 21 | Commerce & eCommerce | 275 | 1.65 |
| 2 | Data & Database | 1,368 | 8.21 | 22 | Devices & IoT | 271 | 1.63 |
| 3 | Business & Enterprise | 1,324 | 7.94 | 23 | Medical | 243 | 1.46 |
| 4 | Finance & Payments | 1,237 | 7.42 | 24 | Visual Recognition & Image Processing | 237 | 1.42 |
| 5 | Advertising & Marketing | 983 | 5.90 | 25 | Storage & File Management | 212 | 1.27 |
| 6 | AI & Machine Learning | 701 | 4.21 | 26 | Events | 211 | 1.27 |
| Configuration | Value |
|---|---|
| Optimizer | AdamW |
| Adam coefficients | , |
| Peak learning rate | |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Learning-rate schedule | Cosine decay to |
| Parameter | Value |
|---|---|
| Temperature | 0.7 |
| Top- | 0.9 |
| Top- | 20 |
| Maximum generated tokens per response | 32,768 |
| Maximum context length | 262,144 |
| Parameter | Value |
|---|---|
| Maximum position embeddings | 262,144 |
| RoPE base ( ) | 1,000,000 |
| RoPE scaling type | YaRN |
| Scaling factor | 6.4 |
| Original maximum position embeddings | 40,960 |
| Benchmark | Evaluation set | Reported metrics | Rollouts/task |
|---|---|---|---|
| WildClaw | Full set | Score, Pass@1 | 1 |
| Claw-Eval | Full set | Score, Passˆ3 | 3 |
| GDPval | 201 text-only tasks | Rubric score | 1 |
| APEX | Full set | Score, Pass@1 | 1 |
| ALE | Full set | Score, Pass@1 | 1 |
| Benchmark | Harness | Wall-clock limit | Interaction limit |
|---|---|---|---|
| WildClaw | OpenClaw | 24,000 s | Not set |
| Claw-Eval | OpenClaw | 3,600 s | 50 turns |
| GDPval | Claude Code | 14,400 s | 300 turns |
| APEX | ReAct Toolbelt | 14,400 s | 250 steps |
| ALE | Claude Code | 28,800 s | Unlimited |
| Row | Task ID | Occupation |
|---|---|---|
| 12 | 38889c3b-e3d4-49c8-816a-3cc8e5313aba | Audio and Video Technicians |
| 13 | ff85ee58-bc9f-4aa2-806d-87edeabb1b81 | Audio and Video Technicians |
| 14 | 4b894ae3-1f23-4560-b13d-07ed1132074e | Audio and Video Technicians |
| 56 | e222075d-5d62-4757-ae3c-e34b0846583b | Film and Video Editors |
| 57 | c94452e4-39cd-4846-b73a-ab75933d1ad7 | Film and Video Editors |
| 58 | 75401f7c-396d-406d-b08e-938874ad1045 | Film and Video Editors |
| Category | Criteria | Share | Base | SFT | |
|---|---|---|---|---|---|
| Visual / layout quality | 371 | 4.2% | 74.1% | 87.1% | +13 pp |
| Numerical claims | 523 | 5.9% | 64.6% | 72.8% | +8.2 pp |
| Cross-source reconciliation | 103 | 1.2% | 86.4% | 92.2% | +5.8 pp |
| Reasoning / argumentation | 200 | 2.2% | 82.0% | 86.5% | +4.5 pp |
| Content coverage | 7,718 | 86.6% | 77.4% | 83.4% | +6.0 pp |
| Group | Turns per trajectory | Trajectories |
|---|---|---|
| G1 | 0–20 | 1,000 |
| G2 | 21–40 | 1,000 |
| G3 | 41–60 | 1,000 |
| Group | Productive calls | Skill discovery | Tool calls | Distinct actions | Tool types used |
|---|---|---|---|---|---|
| G1 | 11.7 | 21.0% | 17.7 | 17.0 | 3.3 |
| G2 | 23.1 | 21.9% | 38.3 | 33.0 | 3.6 |
| G3 | 24.0 | 33.3% | 42.0 | 38.3 | 4.0 |
| Variant | Selection rule | Trajectories |
|---|---|---|
| LLM only | 1,000 | |
| Programmatic only | 1,000 | |
| Hybrid | 1,000 |