Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal.
Figures & tables
Figure 1: Overview of the GraphForge pipeline. Starting from an O*NET-derived seed, an agent assembles a workspace of real files with hidden roles, and a model builds an evidence graph that is compiled into a task specification whose rubric criteria are anchored to graph nodes. An initial teacher rollout supports a one-step revision of the task specification, and the final trajectory is scored by an evidence-anchored judge before admission.
Dataset
Domain
Task environment
Verification
Scale
General-Domain Pipelines
AgentSynth
computer use
desktop VM
per-step execution check
6K tasks
TaskCraft
tool use
web and document tools
golden answer
36K tasks
SWE-smith
software eng.
code repositories
fail-to-pass tests
50K tasks
CLI-Universe
terminal
docker environments
fail-to-pass tests
6K trajs
Working-Agent Pipelines
Table 1: Comparison of training-data synthesis pipelines for agents. The upper block lists general-domain pipelines, the middle block lists working-agent pipelines, and the bottom row shows our pipeline. Scale reports the number of tasks or trajectories produced by each pipeline.
Figure 2: Diversity of the SFT corpus across occupational sectors, execution patterns, and input file families.
Figure 3: Distribution of assistant steps and total tokenized length in the 2,169-example SFT corpus.
Model
GDPVal
Workspace-Bench-Lite
SpreadsheetBench II
OpenHands
Codex
Claude Code
Codex
Claude Code
Codex
Frontier Models
Claude Opus 5
1774.1
1753.1
70.1
68.9
33.6
—
GPT-5.6-sol
1687.1
1710.8
—
60.5
—
32.7
Qwen3.8-Max
1719.0
1771.0
67.4
66.6
34.9
34.9
GLM-5.3
1667.0
1543.7
67.7
61.4
32.1
31.5
Table 2: Overall comparison on working agent benchmarks. GDPVal reports Elo under the OpenHands and Codex scaffolds, with each (model, scaffold) pair fitted as a separate node and the scale anchored at GLM-5.3 (OpenHands) = 1667. Workspace-Bench-Lite reports micro scores and SpreadsheetBench II reports execution accuracy, both under the Claude Code and Codex scaffolds. GDPVal bootstrap confidence intervals are reported in Appendix C .
Model
GDPVal
Workspace-Bench-Lite
SpreadsheetBench II
OpenHands
Codex
Claude Code
Codex
Claude Code
Codex
Our Models
Qwen3.6-35B-A3B (reference)
1260.6
1283.0
55.9
53.4
2.8
4.7
35B SFT (GraphForge)
1362.3
1384.4
59.7
60.0
19.3
18.7
RFT Variants
+ RFT
1369.5 (+7.2)
1395.4 (+11.0)
63.7 (+4.0)
64.0 (+4.0)
20.3 (+1.0)
19.6 (+0.9)
Table 3: RFT ablation on top of the SFT model. Parentheses report the change from the SFT model. GDPVal values are SFT-anchored Elo from direct paired comparisons with SFT. GDPVal bootstrap confidence intervals are reported in Appendix D .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
SFT
RFT arms
Initialization
Qwen3.6 base (27B or 35B-A3B)
35B-A3B SFT checkpoint
Training examples
2,169 trajectories
462 matched trajectories per arm
Data admission
one-step revision, Qi>0.90
best-of-4, Qi>0.95 , behavior-clean
Epochs
3
1
Optimizer
Muon
Muon
Peak learning rate
2×10−5
1×10−6
Appendix
Table 5: SFT and RFT training configurations. The anchored, unanchored, and random RFT arms use the same 462 query IDs, candidate pools, behavior filter, and optimization budget.
Model
OpenHands Elo [95% CI]
Codex Elo [95% CI]
Claude Opus 5
1774.1 [1750.2, 1798.3]
1753.1 [1697.4, 1815.7]
GPT-5.6-sol
1687.1 [1663.4, 1710.7]
1710.8 [1657.8, 1769.3]
Qwen3.8-Max
1719.0 [1694.9, 1741.8]
1771.0 [1714.6, 1831.9]
GLM-5.3
1667.0 [1667.0, 1667.0]
1543.7 [1491.6, 1595.8]
Kimi-K3
1615.5 [1592.6, 1638.4]
1664.4 [1613.9, 1717.2]
DeepSeek-V4-Pro
1531.5 [1507.5, 1555.4]
1576.7 [1527.0, 1628.2]
Appendix
Table 6: GDPVal Elo with 95% bootstrap confidence intervals. OpenHands and Codex results are separate (model, scaffold) nodes.
Model
OpenHands Elo [95% CI]
Codex Elo [95% CI]
35B SFT
1362.3 (fixed)
1384.4 (fixed)
+ RFT
1369.5 [1319.1, 1416.5]
1395.4 [1349.5, 1440.1]
+ RFT (unanchored)
1396.4 [1348.0, 1446.3]
1409.7 [1363.8, 1458.1]
+ RFT (random-of-4)
1353.3 [1306.3, 1401.9]
1374.9 [1330.3, 1420.8]
Appendix
Table 7: RFT direct-comparison Elo with conditional 95% bootstrap confidence intervals.
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows. ARBIGRAPH represents each task as a natural-language problem with an executable Python solver, and composes tasks through typed intermediate states, instantiated here as scalar and list values. This design enables controllable task graphs whose length, dependency structure, distractor count, and value type can be varied while preserving exact automatic verification. We instantiate ARBIGRAPH with math, GSM-style word-problems, and Python-tracing task categories, and evaluate a Qwen3.5-27B tool-assisted agent across four topologies. The results show high accuracy on isolated tasks but substantial degradation on more complex dependent tasks: accuracy drops by up to 33.3% on branching chains of dependent math tasks. This shows that ARBIGRAPH exposes failures that are not visible from single-task evaluation alone. Our code, generated datasets, and evaluation results are available at https://github.com/pavelgolikov/ArbiGraph.git
Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko +1
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.
Xinyu Zhu, Fenyi Liu, Yuzhu Cai +4
Shanghai Jiao Tong University · Endless Frontier · Beihang University
Scaling executable agent training data is bottlenecked by substrate-first methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual expansion of the substrate, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate convenience rather than real-world demand. We introduce NexForge, a requirement-first framework that compiles free-form capability requirements into executable agent training data. NexForge first performs research-based demand discovery to identify representative task forms, realistic scenarios, and their relative prevalence. It then applies distribution-aware task compilation and automatically retrieves or constructs the files, repositories, dependencies, and runtime configurations required to materialize each task, followed by teacher rollout collection and trajectory distillation. The same pipeline, without any domain-specific infrastructure, produces 3,600 terminal tasks and 2,000 office tasks, improving Qwen3.5-35B-A3B Base from 22.5% to 52.0% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval; scaling to 43.2K terminal tasks reaches 58.4%, surpassing Claude Opus 4.6. Scaled further, NexForge-synthesized data contributes to the training of Nex-N2, a family of publicly available agent models that lift Qwen3.5-35B-A3B to 75.3% on Terminal-Bench 2.1 and to 1585 Elo on GDPval -- achieving state-of-the-art open-source performance and surpassing several frontier proprietary systems. Nex-N2 models are available at https://nex.sii.edu.cn/
Jiarong Zhao, Zhikai Lei, Zhiheng Xi +5
1East China Normal University · 2Shanghai Qiji Zhifeng Co., Ltd