WorkGenesis: Building the Worlds That Teach Agents to Work
Organizations: Shanghai Jiao Tong University · Endless Frontier · Beihang University
Abstract
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.
Figures & tables
| GDPvalAA-v2 | APEX-Agents-AA | JobBench | |||||
| Model | Params | Rubric | Elo | pass@1 | Main | Easy | AVG |
| Frontier Models | |||||||
| Qwen3.8-Max | 2.4T-A95B | 91.53 | 1739 | 38.9 | 52.57 | 82.81 | 51.14 |
| Kimi-K3 | 2.8T-A104B | 91.64 | 1687 | 35.5 | 43.81 | 80.72 | 46.22 |
| GPT-5.5 | – | 90.50 | 1491 | 29.9 | 30.10 | 78.22 | 36.52 |
| GLM-5.2 | 753B-A40B | 90.34 | 1510 | 29.7 | 33.79 | 75.49 | 37.80 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Sector | Occupation | Sector | Occupation |
| FIN | Customer Service Representatives | FIN | Financial Managers |
| FIN | Financial and Investment Analysts | FIN | Personal Financial Advisors |
| FIN | Securities, Commodities, and Financial Services Sales Agents | GOV | Administrative Services Managers |
| GOV | Child, Family, and School Social Workers | GOV | Compliance Officers |
| GOV | Court, Municipal, and License Clerks | GOV | First-Line Supervisors of Police and Detectives |
| GOV | Public Safety Telecommunicators | GOV | Recreation Workers |
| Model and initialization | |
| Base model | Qwen3.6-35B-A3B (sparse MoE) |
| Training data and mixture | |
| Update steps | 1,537 |
| Packing bin | 256,000 tokens (16k per GPU over CP) |
| Sampling | shuffled and length-balanced per rollout |
| Optimization | |