RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
Organizations: University of Illinois Urbana-Champaign · National Center for Supercomputing Applications
Abstract
Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.
Figures & tables
Appendix figures & tables50 assets
Supplementary material from the paper’s appendix.
Appendix
| Tool | Function |
|---|---|
| workspace_bash | Bash with the episode workspace as the working directory, on the node that runs the harness, with no GPU passed through. Cloning, installs, edits, inspection. Timeout default 900 s, maximum 3,600 s. Returns stdout, stderr, exit code, and duration. |
| write_file | Write a UTF-8 file under the workspace or evidence roots, up to 200,000 characters, overwriting an existing file at that path. |
| edit_file | Replace one exact string in an existing file. The match must be unique unless replace_all is set. A bare repo-relative path is located by suffix search when it resolves to exactly one file. |
| update_plan | Replace the agent’s ordered checklist of steps with statuses pending , in_progress , or completed . At most one step may be in_progress . |
| fetch_url | Fetch a public http(s) URL and reduce HTML to text, default 24,000 characters. |
| list_partitions | Run sinfo and return each Slurm partition’s node counts by state, walltime limit, and GPU resources, plus the profile default. Read-only. |
| Field | Content |
|---|---|
| paper_id | The paper’s arXiv ID. |
| claim | The claim the agent worked toward, in the agent’s own words. |
| what_ran | What was executed, from the environment build to the run. |
| scoring_command | One command that reproduces the measured number from a clean state. Empty when nothing ran. |
| measurements | Array of measured metrics. Each entry records the metric name with units, the observed value, the paper’s reference value or null , the scope the value was measured over, and an array of citations into evidence/ supporting the observed value. |
| agent_assessment | One of reproduced , partial , not_reproduced , could_not_run . |
| Setting | DeepSeek-V4-Flash | Qwen3.6-27B | MiniMax-M2.7 | Muse Spark 1.2 |
|---|---|---|---|---|
| Served checkpoint | deepseek-ai/DeepSeek-V4-Flash-0731 | Qwen/Qwen3.6-27B-FP8 | cyankiwi/MiniMax-M2.7-AWQ-4bit | muse-spark-1.2-contributor |
| Advertised name | same as checkpoint | same as checkpoint | MiniMaxAI/MiniMax-M2.7 | same as checkpoint |
| Precision | native INT8/FP8 | FP8 | 4-bit AWQ (W4A16) | vendor-served |
| KV cache dtype | fp8 | fp8 | not set | vendor-served |
| GPUs per replica | 2 | 1 | 2 | vendor-served |
| Tensor parallel | 2 | 1 | 2 | vendor-served |
| Rate ($/M) | Tokens (M) | Cost ($) | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Agent | Prompt | Cached | Completion | Prompt | Completion | Tokens | GPU | Total | Per run |
| DeepSeek-V4-Flash | 0.14 | 0.0028 | 0.28 | 1,077.9 | 11.11 | 15.04 | 1,337.64 | 1,352.68 | 13.66 |
| Qwen3.6-27B | 0.45 | – | 2.70 | 489.1 | 6.24 | 236.93 | 797.63 | 1,034.56 | 10.45 |
| MiniMax-M2.7 | 0.30 | 0.06 | 1.20 | 573.8 | 6.62 | 82.21 | 528.61 | 610.82 | 6.17 |
| Muse Spark 1.2 | 1.25 | 0.15 | 4.25 | 1,494.6 | 33.53 | 724.91 | 1,112.34 | 1,837.24 | 18.56 |
| All agents | 3,635.4 | 57.50 | 1,059.08 | 3,776.22 | 4,835.30 | 12.21 | |||
| By tier | By band | |||||
|---|---|---|---|---|---|---|
| Agent | Tier or measure | Reproduced | Mean score | 8 h | 32 h | 96 h |
| DeepSeek-V4-Flash | Run | 14/34 (41%) | 5.29 | |||
| Retrain | 9/33 (27%) | 5.45 | ||||
| Reimplement | 4/33 (12%) | 4.64 | ||||
| All tiers | 27/100 (27%) | 5.13 | ||||
| Mean score | 6.00 | 4.10 | 3.33 | |||
| Tool | Function and limits |
|---|---|
| list_run_files | Lists files and directories in the run bundle, optionally recursive. At most 200 entries, flagged when truncated. Prunes the .git , pycache , .venv , and node_modules subtrees. |
| read_run_file | Reads one text file by relative path. Default cap 40,000 characters, maximum 200,000, flagged when truncated. |
| bash | Runs one shell command with the run bundle as the working directory. Timeout 60 s, after which the whole process group is killed. Returns stdout, stderr, and the exit code. |
| write_run_file | Writes a new text file of up to 200,000 characters into the bundle, such as a re-scoring script to run with bash and cite. Refuses to overwrite an existing file, so agent artifacts stay intact. |
| Score | Band | Conditions | Verdict |
|---|---|---|---|
| 10 | Faithful, exact | Right experiment executed, C1 criterion met, no flags, correct protocol, split, and scale; a stochastic metric seed-fixed or averaged. | reproduced |
| 9 | Faithful, single draw | As 10, with one unseeded draw of a stochastic metric that still lands inside the bar. | reproduced |
| 8 | Minor caveats | C1 criterion met with sound provenance; low-severity flags only, or one medium flag whose evidence shows the deviation does not change the measured quantity. The bar is never loosened at grade time. | reproduced |
| 7 | Near-reproduction | The authors’ own pipeline, clean provenance, honest, and the result lands just outside the bar, or inside it with an unresolved protocol deviation, or all but one arm of a multi-part claim is met. | partial |
| 6 | Clear partial | The right quantity executed with sound provenance, and the result clearly misses the bar, or the pinned bar is mis-specified and the agent reproduced the reproducible sibling quantity. | partial |
| 5 | Honest off-target | No authors’ evaluation pipeline existed, so the agent reconstructed the protocol in good faith, ran the whole evaluation, disclosed the result, and it diverged from the target. | not_reproduced |
| Tier | Selected | 0–8 | 8–32 | 32–96 | 96 | Flagged | Revised | H100-h |
|---|---|---|---|---|---|---|---|---|
| Run ( Easy ) | 67 | 13 | 35 | 14 | 5 | 14 | 6 | 1,961 |
| Retrain ( Medium ) | 67 | 13 | 39 | 11 | 4 | 22 | 3 | 1,754 |
| Reimplement ( Hard ) | 66 | 24 | 23 | 11 | 8 | 14 | 4 | 2,190 |
| Total | 200 | 50 | 97 | 36 | 17 | 50 | 13 | 5,905 |
| Evaluation split | Development split | |||||
| Tier | Papers | 0–8 | 8–32 | 32–96 | Audited H100-h | Papers |
| Run | 34 | 27 | 7 | 0 | 142.71 | 5 |
| Retrain | 33 | 16 | 13 | 4 | 380.91 | 4 |
| Reimplement | 33 | 16 | 9 | 8 | 630.18 | 5 |
| All | 100 | 59 | 29 | 12 | 1,153.8 | 14 |
| Field | Type | Description |
|---|---|---|
| Top-level fields | ||
| custom_id | string | arXiv ID, the primary key. |
| central_claim | string | The paper’s central empirical claim in one sentence. |
| claim_evidence | string | Quoted paper text grounding the claim. |
| mre_config | string | Prose recipe for the cheapest configuration that tests the claim, withheld from the agent. |
| agent_task | string | Classifier-written reproduction instruction, released with the row and shown to neither the agent nor the auditor. |
| arXiv | Tier | Band | H100-h | Metric | Bar | Target value |
|---|---|---|---|---|---|---|
| 2502.05795 | Run | 0–8 | 1.00 | Validation perplexity… | point estimate | 25.76 |
| 2502.06684 | Run | 0–8 | 0.01 | Mean accuracy (%) on 10… | point estimate | 77.0 +/- 1.8% |
| 2503.18430 | Run | 0–8 | 0.64 | Average Precision (AP) | point estimate | 46.3% |
| 2503.23035 | Run | 0–8 | 0.25 | PSNR (dB) | point estimate | 27.69 |
| 2504.20571 | Run | 0–8 | 0.16 | MATH500 accuracy (%) | point estimate | 73.6% |
| 2505.10978 | Run | 0–8 | 2.00 | Success rate (%) | point estimate | 86.7% |
| arXiv | Tier | Band | H100-h | Metric | Bar | Target value |
|---|---|---|---|---|---|---|
| 2505.18513 | Run | 0–8 | 0.01 | LDS (Linear Datamodeling… | point estimate | 21.11 |
| 2506.09045 | Run | 0–8 | 0.01 | end-to-end latency speedup | point estimate | 2.68x |
| 2507.02546 | Run | 0–8 | 0.01 | Metric point map relative… | threshold | 4.44% |
| 2510.21323 | Run | 0–8 | 0.01 | Unified concept extraction… | point estimate | Unified concept set shared… |
| 2511.07099 | Run | 0–8 | 0.01 | Speaker similarity score… | point estimate | 0.113 |
| 2502.06067 | Retrain | 0–8 | 2.00 | 95% CI coverage rate | threshold | 1.0 (100%) |
| Mode | DeepSeek-V4-Flash | Qwen3.6-27B | MiniMax-M2.7 | Muse Spark 1.2 | All |
|---|---|---|---|---|---|
| Reproduced | 27 | 13 | 10 | 22 | 72 |
| Ran, outside tolerance | 26 | 16 | 7 | 27 | 76 |
| Reimplemented but did not check the result | 13 | 16 | 27 | 7 | 63 |
| Build and dependency failures | 1 | 8 | 11 | 4 | 24 |
| Wrong artifact measured | 6 | 4 | 13 | 15 | 38 |
| Wrong experiment | 9 | 11 | 16 | 10 | 46 |
| Mode | Definition | Evidence signature | Rubric anchor |
|---|---|---|---|
| reproduced-clean Reproduced | The agent runs the pinned experiment, using the released artifacts the tier provides, and the graded number lands inside the pinned bar. | A number the run computed, traced to a command in the transcript that ran the experiment on the released artifacts the tier provides. | 8–10 |
| near-miss-partial Ran, outside tolerance | The agent runs the right pipeline and measures a number that misses the bar. | A completed run of the pinned quantity with clean provenance, and a measured value on the wrong side of the bar. | 6–7 |
| reimplement-without-validating Reimplemented but did not check the result | The agent writes its own version of the method, or part of it, and never checks any stage against a reference it did not produce. | A component written from scratch in the run’s own code, and a reference number the paper supplies, most often its baseline in the same table as the target, that the transcript never compares against. | 4 |
| environment-fights Build and dependency failures | Most of the run goes to build and dependency failures in the released stack, and the experiment never produces a valid number. | Repeated install, compile, and binary-compatibility failures against the authors’ pinned versions, with the same packages failing across runs on the same host. | 2 |
| artifact-provenance-mismatch Wrong artifact measured | The agent measures a different checkpoint, dataset, split, or scale from the one the dataset row names. | A downloaded artifact whose model card, README, or example disagrees with the dataset row, and a measurement taken on it. The wrong_split_scale_dataset flag marks the same break. | 0 with a high flag |
| scope-substitution Wrong experiment | The agent runs a different experiment from the one the pinned claim is about. | A round where the agent names the pinned arm and picks a different one, and a graded number from the substitute. | 0 or 2 |