AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
Organizations: HKUST · SLAI · NTU
Abstract
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human-in-the-loop collaboration. Automating task creation would let data production scale with compute rather than with expert headcount, would extend to more domains, and would enable a key step in recursive self-improvement (RSI). Current evaluations of an agent's ability to write such tasks measure how a model performs after training on what the agent produced. That does not match common practice in the data industry, where data is delivered sample by sample and each sample is accepted against a set of criteria rather than put straight into training. No existing evaluation asks whether an individual task meets the acceptance criteria of a data pipeline. We therefore introduce AutoDataBench. Given an original benchmark task and a record of the target model attempting it, an agent must write a new task for the same suite that meets practical acceptance standards on validity, novelty, difficulty and behavioural coverage. Across three benchmarks of executable agent tasks, no agent we evaluate scores above 20 out of 100 at the default time budget of 45 minutes. Giving the strongest agent four times as long improves its score substantially, while the cost of one usable task stays almost unchanged. Current agents can write training tasks of the required quality, but not efficiently. AutoDataBench provides a direct measure of an agent's capacity for autonomous data synthesis: one artifact at a time, judged against the criteria a production pipeline would apply, and without a training run. Code and data are available at https://github.com/StarDewXXX/AutoDataBench.
Figures & tables
| Weakness- | Per-artifact | Difficulty | No training | |||
|---|---|---|---|---|---|---|
| Benchmark | Artifact delivered | targeted | verdict | in a band | run | Domains covered |
| AutoBencher | evaluation items | ✗ | ✗ | ✓ | knowledge, math, multilinguality, safety | |
| BenchAgents | evaluation items | ✗ | ✗ | ✓ | planning, constraints, causal reasoning; text and vision | |
| InnovatorBench | research artifacts, including constructed data | ✗ | ✓ | ✗ | ✗ | LLM research |
| PostTrainBench | a trained checkpoint | ✗ | ✗ | ✗ | ✗ | math, science, code, tool use, writing, health |
| RSIBench-Data | supervision for a fixed SFT interface | ✓ | ✗ | ✗ | ✗ | software, terminal, science QA, math |
| author agent | In band | Gate passed in band | Quality | Score |
|---|---|---|---|---|
| kimi-k3 | 21.3% | 90.0% | 0.967 | 0.184 |
| gpt-5.6-sol | 21.3% | 100.0% | 0.833 | 0.177 |
| qwen3.8-max | 14.6% | 100.0% | 0.952 | 0.139 |
| glm-5.3 | 21.7% | 90.0% | 0.700 | 0.130 |
| deepseek-v4-pro | 25.0% | 41.7% | 0.931 | 0.097 |
| per episode | per usable delivery | |||
|---|---|---|---|---|
| author agent | minutes | $ | minutes | $ |
| deepseek-v4-pro | 42.7 | 0.50 | 410 | 4.81 |
| qwen3.8-max | 41.2 | 1.53 | 283 | 10.49 |
| kimi-k3 | 45.7 | 2.23 | 244 | 11.88 |
| glm-5.3 | 48.4 | 3.50 | 258 | 18.69 |
| gpt-5.6-sol | 42.0 | 5.77 | 202 | 27.69 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| domain | task |
| Terminal-Bench (commit 624df06 ) | |
| Hardware | retro-console-soc |
| ML | batched-eval-parity |
| Media | satb-audio-transcription |
| Operations | medical-claims-processing |
| Science | roy-polymorph-cn |
| The trade-off that frames the rest | |
|---|---|
| C1 | Proximity to the original keeps its modes reachable and invites the surface-swap gate; distance lowers that risk and removes the modes along with the structure they depended on |
| C2 | Keeping the original’s scene and tooling while adding a constraint it lacked satisfies both sides, occurs in the data, and is never adopted as a default |
| Patterns that cost the pass rate | |
| C3 | Incompleteness in the original specification is read as a defect and written out, though it was the source of the task’s discrimination; this costs coverage as well, since a judgement that is spelled out is neither failed nor recorded |
| C4 | A pass rate is extrapolated from one sample, one large one-directional edit follows, and the result is delivered without re-testing |
| C5 | Whether the difficulty-calibration loop can close is decided by the ratio of one target-model attempt to the time budget, not by the agent’s method |