TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development
Organizations: Carnegie Mellon University
Abstract
Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4{,}465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, and the labeling models, together with an open-source toolkit that turns a run from any command-line agent into a TraceML trajectory and reads it against the human cohorts.
Figures & tables
| Benchmark / Dataset | Score Traj | Code Traj | Human Traj | # Envs. | Horizon |
|---|---|---|---|---|---|
| MLE-bench ( Chan et al., 2024 ) | ✗ | ✗ | ⚫ | 75 | 24h |
| MLAgentBench ( Huang et al., 2024 ) | ✗ | ⚫ | ✗ | 13 | – |
| AIRA ( Toledo et al., 2025 ) | ⚫ | ⚫ | ✗ | 22 | 24h |
| RE-Bench ( Wijk et al., 2025 ) | ✓ | ⚫ | ✓ | 7 | 8h |
| HCAST ( Rein et al., 2025 ) | ✗ | ✗ | ✓ | 189 | 1m–8h |
| SciCode ( Tian et al., 2024 ) | ✗ | ✗ | ✗ | 80 / 338 | – |
| Subset | Comps | Trajectories | Snapshots | Avg snap/traj | Avg lines/snap |
| Human (public Kaggle notebook histories) | |||||
| Grandmaster | 114 | 423 | 16,973 | 40.1 | 555 |
| Master | 121 | 649 | 24,138 | 37.2 | 601 |
| Expert | 130 | 1,386 | 50,663 | 36.6 | 569 |
| Contributor | 133 | 1,932 | 54,619 | 28.3 | 494 |
| Other / Unknown | 40 | 75 | 3,090 | 41.2 | 295 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Outcome | Kernels | Character of the set |
|---|---|---|
| Retained | 4,465 | 47.3% medalled |
| Removed: no in-window version | 382 | 99.2% published entirely after the deadline, median 299 days late |
| Removed: failed content conditions | 201 | Too few versions, too short a span, or no score |
| Removed (all) | 583 | 43.3% medalled |
| Reuse measure | Full corpus (4,465) | Paired sample (430) |
|---|---|---|
| Has a fork parent | 8.2% | 8.6% |
| Near-duplicate code in version | 23.7% | 21.0% |
| Near-duplicate code in a majority of versions | — | 7.9% |
| Near-duplicate code across the whole trajectory | — | 3.3% |
| Human | Codex | MLEvolve | |
|---|---|---|---|
| Median versions per trajectory | 20 | 18 | 5 |
| Versions carrying a score | 45% | 99% | 99% |
| Saved versions that are syntactically valid code | 97.9% | — | — |
| MLEvolve | |||||
| Scope | Comps | Human traj. | Codex runs | runs | branches |
| Corpus (§ 3.1 ) | 134 | 4,465 | — | — | — |
| Paired subset (§ 3.2 ) | 7 | 430 | 11 | 13 | 189 |
| Twelve-hour scope (§ 4 ) | 7 | 430 | 10 | 3 | 107 |
| Harness arms (§ 5.1 ) | 7 | — | 30 | — | — |
| Competition | Track | Transitions | Pivot rate |
| gquest | agent | 1,017 | 2.1% |
| gquest | llm_v3 | 226 | 2.2% |
| ranzcr | agent | 166 | 4.8% |
| ranzcr | llm_v3 | 102 | 2.0% |
| hms | agent | 97 | 2.1% |
| aes2 | llm_v3 | 64 | 1.6% |
| Transition | Actions | Intents | Magn. | Score |
|---|---|---|---|---|
| model, training, housekeeping, infra | optimization + debugging | micro | (improving) | |
| data, augmentation, training, model | exploration + optimization | minor | unscored | |
| data, training, validation, housekeeping | debugging + optimization | micro | (trajectory best) | |
| training, model, infra | optimization | micro | (regressing) |
| Annotation level (#tags) | Self-consistency | Cross-model | Teacher Student |
|---|---|---|---|
| State coarse (8) | / J= | / J= | |
| Action coarse (10) | J= | J= | |
| Intent (6 classes) | |||
| Magnitude (4 levels) |
| Cohort | 6-class agreement | 3-class collapse | |
|---|---|---|---|
| Human | 40 | 80.0% | 82.5% |
| Codex | 40 | 85.0% | 87.5% |
| MLEvolve | 20 | 75.0% | 75.0% |
| Overall | 100 | 81% | 83% ( ) |
| Gap (human agent) | Orig. | Scored-only | Agents filt. | New runs | 95% cluster CI |
|---|---|---|---|---|---|
| Debugging intent, vs Codex | |||||
| Debugging intent, vs MLEvolve | — | ||||
| Ensemble action mass, vs Codex | |||||
| Model train. mass, vs MLEvolve | — | ||||
| Action JSD (bits), vs Codex | |||||
| Action JSD (bits), vs MLEvolve | — |
| Coarse action | Top-40% humans | All humans | Codex | MLEvolve |
|---|---|---|---|---|
| data | 8.0 | 9.1 | 8.5 | 9.0 |
| features | 9.1 | 10.0 | 5.6 | 11.9 |
| augmentation | 0.4 | 0.4 | 0.2 | 0.4 |
| model | 10.7 | 10.5 | 8.2 | 20.4 |
| training | 11.4 | 11.4 | 8.6 | 18.0 |
| ensemble | 9.8 | 6.3 | 19.7 | 6.0 |
| Cohort | Quintile | data | feat. | aug. | model | train. | ens. | valid. | infer. | infra | house. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Top-40% humans | Q1 | 9.4 | 8.7 | 0.2 | 9.8 | 10.3 | 7.3 | 4.1 | 9.6 | 15.1 | 25.4 |
| Q2 | 7.8 | 8.7 | 0.6 | 10.4 | 13.5 | 9.1 | 4.9 | 9.2 | 13.5 | 22.4 | |
| Q3 | 7.7 | 9.0 | 0.3 | 10.8 | 10.6 | 10.1 | 5.9 | 10.3 | 13.0 | 22.2 | |
| Q4 | 6.6 | 9.1 | 0.4 | 9.4 | 11.1 | 10.5 | 5.0 | 10.8 | 12.5 | 24.5 | |
| Q5 | 6.5 | 9.3 | 0.3 | 10.6 | 11.3 | 11.3 | 5.2 | 10.9 | 12.0 | 22.6 | |
| Codex | Q1 | 8.7 | 5.4 | 0.2 | 6.3 | 8.6 | 22.6 | 9.7 | 23.9 | 11.0 | 3.4 |
| Top humans | Other humans | Codex | MLEvolve | |
| Solution revisit (state signature) | ||||
| % of eligible versions | 9.1% | 7.2% | 0.2% | 0.0% |
| % of trajectories with one | 32.8% | 33.7% | 14.3% | 0.0% |
| eligible versions | 5,838 | 5,648 | 658 | 344 |
| revisits observed | 531 | 406 | 1 | 0 |
| % that beat the version returned to | 78.5% | 63.9% | too small | |
| Feature | One-line description |
|---|---|
| Ensemble timing (3) | |
| has_ens | Indicator for whether any ensemble-blending action appears in the trajectory. |
| ens_late_minus_early | Q5 minus Q1 share of ensemble-blending actions; positive means ensembling concentrates late. |
| pos_first_ens | Normalised position of the first ensemble action; if no ensemble action ever fires. |
| Magnitude and quintile-shift (4) | |
| micro_q1 | Q1 share of transitions labeled magnitude micro (small early edits typical of expert iteration). |
| paired ( , 7 comps) | humans-only ( , 127 comps) | |||
| # | feature | feature | ||
| 1 | fold_averaging | pos_first_ens | ||
| 2 | holdout_split | fold_averaging | ||
| 3 | mode.ens | holdout_split | ||
| 4 | pos_first_ens | blending | ||
| 5 | oof_prediction | change_weights | ||
| Action tag | One-line description |
|---|---|
| Codex bookkeeping-loop markers | |
| change_weights | Re-weight the contributions of existing ensemble members. |
| postprocess_change | Adjust a post-hoc step on predictions (rounding, clipping, calibration). |
| data_loading | Modify how data is read or assembled without changing features or model. |
| add_member | Add a model to the ensemble. |
| MLEvolve mutate-in-place markers | |
| Competition | Metric | Baseline | Harness | Abl-A | Abl-B |
|---|---|---|---|---|---|
| commonlit | RMSE | 0.510 | 0.505 / 0.517 | 0.520 | 0.512 |
| equity | C-index | 0.675 | 0.670 / 0.680 | 0.672 | 0.670 |
| gquest | Spearman | 0.371 | 0.429 | 0.372 | 0.388 |
| aes2 | QWK | 0.771 | 0.817 / 0.808 | 0.806 | 0.796 |
| hms | KL | 1.050 | 0.718 | 0.795 | 1.377 |
| ranzcr | AUC | 0.545 | 0.877 | 0.542 | 0.583 |