cs.LGAug 26, 2026

TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development

Authors: Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang

Organizations: Carnegie Mellon University

Abstract

Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4{,}465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, and the labeling models, together with an open-source toolkit that turns a run from any command-line agent into a TraceML trajectory and reads it against the human cohorts.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TuiML: Machine Learning for AI Agents

    Sep 16, 2026Nilesh Verma, Nick Lim, Albert Bifet +1Artificial Intelligence AgentsModel Context Protocol

  2. ConflictGuide: AutoResearch Improves When Competing Behaviors Are Made Visible

    Sep 30, 2026Binqian Xu, Qiran Zou, Xiangbo Shu +1Code QualityAdversarial Evaluation

  3. Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes

    May 7, 2026Jingjie Ning, Xiaochuan Li, Ji Zeng +2Auto ResearchLoop