cs.AISep 27, 2026

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Authors: Dehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang, Ziyi Chen, Yan Zhang, Qinbo Bai, Mengyuan Chao, +8 more

Organizations: ByteDance Inc., USA · University of Illinois at Chicago

Abstract

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

    Aug 24, 2026Wenhao Wu, Menghao Zhang, Xin Wang +3Large Language Model AgentsSelf-Evolving Skill Libraries

  2. BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

    May 28, 2026Jiahao Huang, Fei Cheng, Junfeng Jiang +2Self-EvolutionLarge Language Model Agents

  3. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

    Jul 7, 2026Andrey Podivilov, Vadim Lomshakov, Sergey Savin +4Coding AgentsFormal Verification