Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collaboration. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent's ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where data is delivered sample by sample and each sample is accepted against a set of criteria rather than put straight into training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task for the same suite that meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. AutoDataBench provides a direct measure of an agent's capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run. Code and data are available at https://github.com/StarDewXXX/AutoDataBench.
Figures & tables
Weakness-
Per-artifact
Difficulty
No training
Benchmark
Artifact delivered
targeted
verdict
in a band
run
Domains covered
AutoBencher
evaluation items
∙
✗
✗
✓
knowledge, math, multilinguality, safety
BenchAgents
evaluation items
✗
∙
✗
✓
planning, constraints, causal reasoning; text and vision
InnovatorBench
research artifacts, including constructed data
✗
✓
✗
✗
LLM research
PostTrainBench
a trained checkpoint
✗
✗
✗
✗
math, science, code, tool use, writing, health
RSIBench-Data
supervision for a fixed SFT interface
✓
✗
✗
✗
software, terminal, science QA, math
Table 1: Benchmarks in which an agent produces data or tasks, and where AutoDataBench differs. Weakness-targeted is whether the artifact must aim at behaviour a designated model has actually been observed to exhibit. Per-artifact verdict is whether the score attaches to one produced artifact rather than to a dataset or a trained checkpoint. Difficulty in a band is whether difficulty is measured on that model and required to fall in a range; AutoBencher measures difficulty on models but maximises it, which is the right objective for an evaluation item and the wrong one for training data. ✓ denotes present, ✗ absent, ∙ partial.
author agent
In band
Gate passed ∣ in band
Quality
Score
kimi-k3
21.3%
90.0%
0.967
0.184
gpt-5.6-sol
21.3%
100.0%
0.833
0.177
qwen3.8-max
14.6%
100.0%
0.952
0.139
glm-5.3
21.7%
90.0%
0.700
0.130
deepseek-v4-pro
25.0%
41.7%
0.931
0.097
Table 2: Main results, means over two episodes per original task. In band is the fraction of deliveries whose target model pass rate fell inside [0.125,0.75] . Gate passed is the fraction of those that also cleared all eight defects. Quality is mean rubric coverage over deliveries that are both in band and ungated. Score is Equation 2 averaged per original task and then across tasks.
per episode
per usable delivery
author agent
minutes
$
minutes
$
deepseek-v4-pro
42.7
0.50
410
4.81
qwen3.8-max
41.2
1.53
283
10.49
kimi-k3
45.7
2.23
244
11.88
glm-5.3
48.4
3.50
258
18.69
gpt-5.6-sol
42.0
5.77
202
27.69
Table 3: What one usable delivery costs, in wall clock and in money. Dollars are US dollars. Usable counts deliveries that were both in band and ungated, out of 48 episodes. Money counts the author agent’s own tokens at September 2026 list prices and excludes the target-model calls it makes while calibrating, the K official attempts and the judge; oracle checks call no model. glm-5.3 is corrected for a cache that never engaged, which is why its measured spend of 746.55doesnotappear:atthemedianhitrateoftheotheragentsonthesameframeworkitwouldhavespent168.20.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
domain
task
Terminal-Bench (commit 624df06 )
Hardware
retro-console-soc
ML
batched-eval-parity
Media
satb-audio-transcription
Operations
medical-claims-processing
Science
roy-polymorph-cn
Appendix
Table 4: The 24 original tasks, with the domain label each suite gives them.
The trade-off that frames the rest
C1
Proximity to the original keeps its modes reachable and invites the surface-swap gate; distance lowers that risk and removes the modes along with the structure they depended on
C2
Keeping the original’s scene and tooling while adding a constraint it lacked satisfies both sides, occurs in the data, and is never adopted as a default
Patterns that cost the pass rate
C3
Incompleteness in the original specification is read as a defect and written out, though it was the source of the task’s discrimination; this costs coverage as well, since a judgement that is spelled out is neither failed nor recorded
C4
A pass rate is extrapolated from one sample, one large one-directional edit follows, and the result is delivered without re-testing
C5
Whether the difficulty-calibration loop can close is decided by the ratio of one target-model attempt to the time budget, not by the agent’s method
Appendix
Table 5: Behaviours recurring across all five author agents, grouped by the score term each one costs. Eight of the fourteen bear on the pass rate, which is the same conclusion Section 3 reaches from the scores.