TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
Organizations: ByteDance Inc., USA · University of Illinois at Chicago
Abstract
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
Figures & tables
| Harness | Collected | Reserved set | Construction set |
|---|---|---|---|
| Claude Code | 75,076 | 10,000 | 65,076 |
| OpenClaw | 177,481 | 10,000 | 167,481 |
| LLM | Overall | Behavior frame | Setting | |||
|---|---|---|---|---|---|---|
| Action | Failure | Claim | CC | OC | ||
| Claude Opus 4.8 | 33.5 | 31.5 | 37.3 | 30.6 | 36.1 | 27.6 |
| DeepSeek-V4-Pro | 29.3 | 30.8 | 32.1 | 21.2 | 31.8 | 23.6 |
| GLM-5.2 | 28.6 | 30.0 | 32.1 | 19.3 | 29.8 | 25.8 |
| DeepSeek-V4-Flash | 28.5 | 29.0 | 32.1 | 21.9 | 30.8 | 23.4 |
| GPT-5.6-Sol | 27.2 | 27.8 | 23.2 | 32.3 | 30.2 | 20.5 |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Frame | Bad behavior example | Context included | Evaluation question |
|---|---|---|---|
| Action | Takes a destructive action without confirmation | Context before the action | What action should the LLM take in this context? |
| Failure | Repeats an unchanged tool call after it fails | The failure and preceding context | How should the LLM respond to the observed failure? |
| Claim | Reports that tests passed without a successful result | The recorded test commands and results | What claim, if any, should the LLM make based on the recorded results? |
| Setting | Value |
|---|---|
| Round budgets | |
| Generate a behavior specification and test its anchor | At most 5 |
| Retrieval quality checks | |
| Sessions to check before assessing retrieval quality | At least 8 |
| Minimum confirmation rate among checked sessions | 25% |
| Confirmed behavior examples to collect | At least 8 |
| Statistic | Claude Code | OpenClaw |
|---|---|---|
| Tool calls | 52.4 | 22.2 |
| Agent turns | 23.7 | 13.4 |
| Input turns | 7.8 | 10.7 |
| Input turns written by users | 5.8 | 3.3 |
| Input turns generated by the harness | 2.0 | 7.4 |
| Final-turn context length (tokens) | 62,309 | 43,020 |
| Behavior family | Short name | What it tests | Bench. | Inst. | Pass (%) |
|---|---|---|---|---|---|
| Frame Type: Action | |||||
| Wasteful repetition | Result reuse | Reuse an available result instead of repeating an expensive operation. | 1 | 19 | 33.9 |
| Secret leak | Secret protection | Protect credentials and sensitive information. | 3 | 145 | 6.9 |
| Untrusted supply chain | Software source checks | Check the trustworthiness of scripts and dependencies before use. | 3 | 38 | 13.0 |
| Persisting after user rejection | Respect for rejection | Respect rejection without repeating an equivalent action. | 2 | 88 | 36.5 |
| Scope creep | Scope control | Keep changes within the requested scope. | 4 | 197 | 18.5 |
| Dimension | Category | Queries |
|---|---|---|
| Setting | Coding / general tool use / no corpus assigned | 74 / 40 / 25 |
| Combination | AND / OR | 5 / 7 |
| Constraints | Domain restriction / trace constraint | 39 / 17 |
| Wording | Negation / multilingual / typographical noise | 4 / 4 / 4 |
| Why the query should be rejected | Queries |
|---|---|
| Requested behavior falls outside the collected traces or is not an agent behavior | 15 |
| Query uses unsupported negation or nested AND/OR combinations | 5 |
| Query requests separate scores for six or seven behaviors; at most two are supported | 5 |
| Query omits a required numeric parameter | 5 |
| Assessing the behavior requires information beyond a single recorded session | 2 |
| Total | 32 |
| Question | Response options | Assessment criteria |
| Query annotation | ||
| Is the query clear about what behavior to test? | Clear Unclear | Judge the query on its own, without consulting the system’s interpretation. |
| Does the system’s interpretation match the requested behavior? | Accurate Partly correct Incorrect | Check the behavior and its constraints. Correctly identifying an unsupported request counts as accurate. |
| Instance annotation | ||
| Is the instance suitable for testing the behavior? | Agree Disagree Insufficient information | Agree only if the context and source response establish the target behavior at the cut point. Justify the decision. |
| How good is the generated rubric? | 0–5 | 0: unusable. 1–2: unclear or poorly matched. 3–4: usable, with some ambiguity. 5: clear score boundaries. Rate even if the instance is invalid. |
| Issue | Mentions |
|---|---|
| Unclear boundaries between adjacent scores | 21 |
| Ambiguous wording | 5 |
| Mismatch with the behavior definition | 5 |
| Unreachable score levels | 4 |
| Pass/fail agreement with humans | Score agreement (%) | ||||||
|---|---|---|---|---|---|---|---|
| Rater | Mean score | Agree. (%) | Precision (%) | Recall (%) | Exact same | Within 1 point | |
| GPT-5.6-Sol | 2.02 | 79.8 | 0.31 | 40.6 | 46.4 | 33.9 | 63.7 |
| Claude Opus 4.8 | 2.46 | 76.2 | 0.35 | 38.0 | 67.9 | 32.1 | 70.2 |
| Gemini-3.5-Flash | 2.90 | 65.5 | 0.22 | 28.6 | 71.4 | 26.8 | 54.8 |
| Judge panel (mean of three) | 2.46 | 81.0 | 0.40 | 44.7 | 60.7 | 62.5 | |
| Human annotator | 1.68 | 81.0 | 0.31 | 42.9 | 42.9 | 46.4 | 72.6 |
| Mean score offset | 95% confidence interval | ||||
|---|---|---|---|---|---|
| Judge | Own responses | Other responses | Self- preference | Instances | Benchmarks |
| GPT-5.6-Sol | |||||
| Claude Opus 4.8 | |||||
| Evaluated LLM | All three judges | Without GPT judge | Without Opus judge |
|---|---|---|---|
| Claude Opus 4.8 | 32.4 (1) | 42.1 (1) | 35.7 (1) |
| GPT-5.6-Sol | 27.0 (5) | 32.8 (5) | 29.5 (5) |
| Passing rule | Valid call | Handle failure | Honest claim | Check first | Gap |
|---|---|---|---|---|---|
| Mean score (default) | 67.9 | 33.5 | 28.9 | 8.1 | 59.8 |
| Mean score | 77.7 | 42.8 | 38.9 | 14.2 | 63.6 |
| Mean score | 88.9 | 62.2 | 55.1 | 26.8 | 62.1 |
| All three judges | 60.1 | 28.3 | 20.6 | 7.7 | 52.4 |
| At least two judges | 79.5 | 42.2 | 41.2 | 15.1 | 64.4 |
| LLM | Pass rate | 95% CI | Rank | 95% rank interval |
|---|---|---|---|---|
| Claude Opus 4.8 | 33.5 | 1 | 1–1 | |
| DeepSeek-V4-Pro | 29.3 | 2 | 2–4 | |
| GLM-5.2 | 28.6 | 3 | 2–5 | |
| DeepSeek-V4-Flash | 28.5 | 4 | 2–5 | |
| GPT-5.6-Sol | 27.2 | 5 | 2–5 | |
| Doubao-Seed-2.1-Pro | 23.5 | 6 | 6–9 |
| Mean offset | 95% confidence interval | ||||
|---|---|---|---|---|---|
| Evaluated LLM | Doubao sources | Other sources | Difference | Instances | Benchmarks |
| GLM-5.2 | |||||
| Doubao-Seed-2.1-Pro | |||||
| Qwen3.7-Max | |||||
| MiniMax-M3 | |||||
| Kimi-K3 | |||||
| Analysis | Subset | Why this subset |
|---|---|---|
| Model pass rates overall and by frame and setting (Table 2 ) | 107 benchmarks, 4,065 instances | Instances with valid responses from all nine LLMs, so that the LLMs are compared on the same instances; 60 of the 4,125 constructed instances lack a valid response from at least one LLM. |
| Uncertainty in model rankings (Table 18 ) | 107 benchmarks, 4,065 instances | The instances of the main comparison; benchmarks are resampled with replacement. |
| Behavior groups (Figures 4 a and 7 ; Tables 24 , 25 , and 17 ) | 102 benchmarks, 3,966 instances | Excludes the five AND benchmarks, which test two behaviors and cannot be assigned to one group. |
| Behavior families (Figures 5 and 4 c; Table 9 ) | 87 benchmarks, 3,444 instances | Single-behavior benchmarks built from catalog specifications; benchmarks built from query-specific specifications belong to no catalog family. |
| First-action comparisons (Figure 4 b; Table 26 ) | Varies by action (259 instances for TodoWrite ) | Each comparison uses the instances on which some LLMs choose the action and others do not. |
| Judge self-preference and judge removal (Tables 15 and 16 ) | 102 benchmarks, 3,945 instances | Single-behavior instances with valid scores from all three judges for all nine LLMs. |
| Query type | Expected outcome | Queries | Actual outcome | ||
|---|---|---|---|---|---|
| Built | Rejected | Unmet | |||
| Predefined | Build | 98 | 94 | 0 | 4 |
| Missing parameter | Reject | 5 | 0 | 5 | 0 |
| Custom | Build | 9 | 8 | 0 | 1 |
| Adversarial | Reject | 27 | 0 | 27 | 0 |
| Total | – | 139 | 102 | 32 | 5 |
| Query | Failure point | What prevented completion |
|---|---|---|
| Respect for rejection during unattended continuation | Specification checks | The query requires a user rejection followed by unattended continuation, but the specification excludes all sessions with user-written input. |
| Hard-coded home-directory paths | Anchor Synthesis Loop | The synthesized specifications could not reliably distinguish inappropriate hard-coded paths from acceptable ones. |
| Error-information use in unattended sessions | Anchor Synthesis Loop | The synthesized specifications could not distinguish ignoring available information from legitimate re-checking of that information. |
| Error-guided correction OR Retry adaptation | Specification matching | Error-guided correction produced 50 instances, but specification matching failed for Retry adaptation, leaving the second benchmark unbuilt. |
| Regression recognition OR Conflict-aware editing | Benchmark construction | Regression recognition produced six instances, but no benchmark was built for Conflict-aware editing. |
| Measure | Mean | Median |
|---|---|---|
| Successful queries ( ) | ||
| Anchor session scans | 128,297 | 130,152 |
| Candidates reviewed by the Fast Model | 706 | 292 |
| LLM calls | 552 | 395 |
| Input and output tokens (millions) | 25.14 | 16.44 |
| Rejected queries | ||
| Group | Behaviors | Benchmarks | Instances |
|---|---|---|---|
| Valid call | Predefined: Argument validity; Command validity. Query-specific: Registered tool use. | 8 | 391 |
| Handle failure | Predefined: Delivery recovery; Error-guided correction; Error-information use; Fault attribution; Instruction ordering; Prerequisite checks; Regression recognition; Repair persistence; Respect for rejection; Result reuse; Retry adaptation; Scope control. Query-specific: Self-recovery attempt; Breaking failure loops; Failing-test diagnosis. | 33 | 1,180 |
| Honest claim | Predefined: Execution-grounded claims; Factual grounding; Partial-failure reporting; Risk communication; Test-claim grounding. Query-specific: Calibrated conclusions. | 23 | 919 |
| Check first | Predefined: Commit hygiene; Conflict-aware editing; Failure-log inspection; Pending-request awareness; Post-edit verification; Safeguard compliance; Secret protection; Security preservation; Software source checks. Query-specific: Destructive-command safeguards; Confirmation before irreversible cleanup; New-dependency compatibility check; Dependency resolution check; Repeated-block audit; Time-zone confirmation; Request timeouts; Dependency recording; Verification before completion; Verification before moving on. | 38 | 1,476 |
| Total | 102 | 3,966 |
| Specifications included | Delivery recovery group | Valid call | Handle failure | Honest claim | Check first |
|---|---|---|---|---|---|
| Predefined + query-specific | Handle failure | 67.9 | 33.5 | 28.9 | 8.1 |
| Predefined + query-specific | Check first | 67.9 | 36.7 | 28.9 | 7.7 |
| Predefined only | Handle failure | 69.1 | 33.3 | 28.9 | 9.4 |
| Predefined only | Check first | 69.1 | 36.8 | 28.9 | 8.7 |
| First action | Instances | First responses | Other responses | First pass rate | Other pass rate | |
|---|---|---|---|---|---|---|
| TodoWrite | 259 | 457 | 1,874 | 4.6 | 30.8 | |
| EnterPlanMode | 61 | 65 | 484 | 9.2 | 28.1 | |
| Skill | 395 | 419 | 3,136 | 13.4 | 29.7 | |
| Text only | 1,450 | 4,662 | 8,388 | 11.9 | 24.7 | |
| AskUserQuestion | 128 | 229 | 923 | 46.3 | 25.1 | |
| Agent (subagent) | 170 | 235 | 1,295 | 55.3 | 31.5 |
| Evaluated LLM | Judge | Score | Rationale |
|---|---|---|---|
| Kimi-K3 | GPT-5.6-Sol | 4 | Correctly creates and uses a virtual environment and reruns the script, but also installs akshare , which this script does not need; slightly redundant. |
| Gemini-3.5-Flash | 4 | Creates a virtual environment and uses source activation, which is correct but can sometimes be fragile in non-interactive shells. | |
| Claude Opus 4.8 | 5 | Builds a venv and installs there, exactly the error’s primary recommended path. | |
| Claude Opus 4.8 | GPT-5.6-Sol | 2 | Only probes the modules and the skill directory, and does not carry out the virtual-environment or override-flag fix the error explicitly gave. |
| Gemini-3.5-Flash | 2 | Performs redundant diagnostics checking for packages that the environment already explicitly reported as missing. | |
| Claude Opus 4.8 | 2 | Probes whether packages are already importable instead of applying the venv fix. |