An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
Figures & tables
Figure 1 : TraceDance’s automated benchmark construction workflow. Given a behavior query and deployment traces, TraceDance selects or synthesizes a behavior specification, retrieves and confirms occurrences, and constructs context instances with a behavior-specific rubric.
Figure 2 : Overview of deployment settings, domains, task types, and undesirable behaviors. Session counts cover the full collection. Domain percentages are relative to each setting’s analysis sample; task-type percentages are relative to the corresponding domain.
Harness
Collected
Reserved set
Construction set
Claude Code
75,076
10,000
65,076
OpenClaw
177,481
10,000
167,481
Table 1 : Session counts by deployment setting.
Figure 3 : System validation and construction workload. (a) Queries meeting (green) or not meeting (pink) their expected build or reject outcome. (b) Mean numbers of session scans, candidates reviewed, and benchmark instances constructed per successful query. (c) Human validation of instances, rubrics, and automated grading.
LLM
Overall
Behavior frame
Setting
Action
Failure
Claim
CC
OC
Claude Opus 4.8
33.5
31.5
37.3
30.6
36.1
27.6
DeepSeek-V4-Pro
29.3
30.8
32.1
21.2
31.8
23.6
GLM-5.2
28.6
30.0
32.1
19.3
29.8
25.8
DeepSeek-V4-Flash
28.5
29.0
32.1
21.9
30.8
23.4
GPT-5.6-Sol
27.2
27.8
23.2
32.3
30.2
20.5
Table 2 : Pass rates (%, higher is better) of nine frontier LLMs on the constructed benchmarks. CC and OC denote Claude Code and OpenClaw, respectively.
Figure 4 : (a) Mean pass rates across benchmarks with different behavior requirements. (b) For each listed first action, bars show its pass rate minus the average pass rate of responses taking other first actions on the same instances. (c) Overall pass rates and pass rates for Error-guided correction .
Figure 5 : Mean pass rates for 28 behavior families represented by benchmarks built from catalog specifications, ordered from highest to lowest. Short names identify the evaluated capabilities. Higher rates indicate more appropriate responses, not more frequent or severe undesirable behavior. Appendix B.3 maps each short name to its behavior family and explains what it tests.
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Frame
Bad behavior example
Context included
Evaluation question
Action
Takes a destructive action without confirmation
Context before the action
What action should the LLM take in this context?
Failure
Repeats an unchanged tool call after it fails
The failure and preceding context
How should the LLM respond to the observed failure?
Claim
Reports that tests passed without a successful result
The recorded test commands and results
What claim, if any, should the LLM make based on the recorded results?
Appendix
Table 3 : The three frames and representative bad behaviors.
Setting
Value
Round budgets
Generate a behavior specification and test its anchor
At most 5
Retrieval quality checks
Sessions to check before assessing retrieval quality
At least 8
Minimum confirmation rate among checked sessions
25%
Confirmed behavior examples to collect
At least 8
Appendix
Table 4 : Limits and sample requirements for benchmark construction.
Table 5 : Prompt template for catalog matching.
Table 6 : Prompt template for candidate confirmation.
Table 7 : Prompt template for response grading.
Statistic
Claude Code
OpenClaw
Tool calls
52.4
22.2
Agent turns
23.7
13.4
Input turns
7.8
10.7
Input turns written by users
5.8
3.3
Input turns generated by the harness
2.0
7.4
Final-turn context length (tokens)
62,309
43,020
Appendix
Table 8 : Statistics of the collected deployment sessions. All values are per-session averages.
Behavior family
Short name
What it tests
Bench.
Inst.
Pass (%)
Frame Type: Action
Wasteful repetition
Result reuse
Reuse an available result instead of repeating an expensive operation.
1
19
33.9
Secret leak
Secret protection
Protect credentials and sensitive information.
3
145
6.9
Untrusted supply chain
Software source checks
Check the trustworthiness of scripts and dependencies before use.
3
38
13.0
Persisting after user rejection
Respect for rejection
Respect rejection without repeating an equivalent action.
2
88
36.5
Scope creep
Scope control
Keep changes within the requested scope.
4
197
18.5
Appendix
Table 9 : The 28 predefined behavior families, their evaluation criteria, and results.
Dimension
Category
Queries
Setting
Coding / general tool use / no corpus assigned
74 / 40 / 25
Combination
AND / OR
5 / 7
Constraints
Domain restriction / trace constraint
39 / 17
Wording
Negation / multilingual / typographical noise
4 / 4 / 4
Appendix
Table 10 : Test queries by deployment setting, behavior combination, constraints, and wording.
Why the query should be rejected
Queries
Requested behavior falls outside the collected traces or is not an agent behavior
15
Query uses unsupported negation or nested AND/OR combinations
5
Query requests separate scores for six or seven behaviors; at most two are supported
5
Query omits a required numeric parameter
5
Assessing the behavior requires information beyond a single recorded session
2
Total
32
Appendix
Table 11 : Types and counts of test queries expected to be rejected.
Question
Response options
Assessment criteria
Query annotation
Is the query clear about what behavior to test?
Clear Unclear
Judge the query on its own, without consulting the system’s interpretation.
Does the system’s interpretation match the requested behavior?
Accurate Partly correct Incorrect
Check the behavior and its constraints. Correctly identifying an unsupported request counts as accurate.
Instance annotation
Is the instance suitable for testing the behavior?
Agree Disagree Insufficient information
Agree only if the context and source response establish the target behavior at the cut point. Justify the decision.
How good is the generated rubric?
0–5
0: unusable. 1–2: unclear or poorly matched. 3–4: usable, with some ambiguity. 5: clear score boundaries. Rate even if the instance is invalid.
Appendix
Table 12 : Questions, response options, and criteria for human annotation.
Issue
Mentions
Unclear boundaries between adjacent scores
21
Ambiguous wording
5
Mismatch with the behavior definition
5
Unreachable score levels
4
Appendix
Table 13 : Rubric issues flagged by annotators; multiple flags are allowed.
Pass/fail agreement with humans
Score agreement (%)
Rater
Mean score
Agree. (%)
κ
Precision (%)
Recall (%)
Exact same
Within 1 point
GPT-5.6-Sol
2.02
79.8
0.31
40.6
46.4
33.9
63.7
Claude Opus 4.8
2.46
76.2
0.35
38.0
67.9
32.1
70.2
Gemini-3.5-Flash
2.90
65.5
0.22
28.6
71.4
26.8
54.8
Judge panel (mean of three)
2.46
81.0
0.40
44.7
60.7
62.5
Human annotator
1.68
81.0
0.31
42.9
42.9
46.4
72.6
Appendix
Table 14 : Mean scores and agreement with human ratings on 84 instances. Each judge and the judge panel are compared with both annotators (168 pairs); the annotators are compared with each other (84 pairs). Scores of at least 4 count as passes, without rounding. The human mean uses all 168 ratings. κ is Cohen’s kappa for pass/fail decisions; precision and recall treat human passes as the positive class. Exact score agreement is omitted for the judge panel’s fractional scores.
Figure 6 : Score distributions for human annotators and the judge panel (%).
Mean score offset
95% confidence interval
Judge
Own responses
Other responses
Self- preference
Instances
Benchmarks
GPT-5.6-Sol
−0.16
−0.38
+0.22
[0.19,0.25]
[0.17,0.27]
Claude Opus 4.8
+0.04
+0.05
−0.01
[−0.03,0.02]
[−0.04,0.03]
Appendix
Table 15 : Relative self-preference in judge scores. Each judge scores 3,945 responses from its own model and 31,560 from other models. Self-preference is the own-minus-other difference in mean score offsets; 95% confidence intervals resample instances or benchmarks.
Evaluated LLM
All three judges
Without GPT judge
Without Opus judge
Claude Opus 4.8
32.4 (1)
42.1 (1)
35.7 (1)
GPT-5.6-Sol
27.0 (5)
32.8 (5)
29.5 (5)
Appendix
Table 16 : Pass rates (%) for Claude Opus 4.8 and GPT-5.6-Sol when grading with all three judges or excluding the GPT or Opus judge. Parentheses show each model’s rank among the nine LLMs. Results use the 102 single-behavior benchmarks.
Passing rule
Valid call
Handle failure
Honest claim
Check first
Gap
Mean score ≥4 (default)
67.9
33.5
28.9
8.1
59.8
Mean score ≥3.5
77.7
42.8
38.9
14.2
63.6
Mean score ≥3
88.9
62.2
55.1
26.8
62.1
All three judges ≥4
60.1
28.3
20.6
7.7
52.4
At least two judges ≥4
79.5
42.2
41.2
15.1
64.4
Appendix
Table 17 : Pass rates (%) of the four behavior groups under alternative passing rules. Gap is Valid call minus Check first in percentage points.
Figure 7 : Pass rates of the four behavior groups under the judge panel and under each judge alone.
LLM
Pass rate
95% CI
Rank
95% rank interval
Claude Opus 4.8
33.5
[28.7,38.2]
1
1–1
DeepSeek-V4-Pro
29.3
[24.8,34.0]
2
2–4
GLM-5.2
28.6
[24.2,33.2]
3
2–5
DeepSeek-V4-Flash
28.5
[24.3,32.8]
4
2–5
GPT-5.6-Sol
27.2
[23.4,31.1]
5
2–5
Doubao-Seed-2.1-Pro
23.5
[19.5,27.7]
6
6–9
Appendix
Table 18 : Overall pass rates (%) with 95% confidence intervals and 95% rank intervals from 5,000 resamples of the 107 benchmarks.
Mean offset
95% confidence interval
Evaluated LLM
Doubao sources
Other sources
Difference
Instances
Benchmarks
GLM-5.2
+1.3
+2.2
−1.0
[−2.9,+1.1]
[−3.2,+1.1]
Doubao-Seed-2.1-Pro
−4.1
−3.2
−0.9
[−2.9,+1.1]
[−3.2,+1.4]
Qwen3.7-Max
−3.6
−2.7
−0.9
[−3.1,+1.3]
[−3.5,+1.7]
MiniMax-M3
−4.1
−3.4
−0.8
[−2.9,+1.3]
[−2.8,+1.3]
Kimi-K3
−4.4
−4.3
−0.1
[−2.3,+1.9]
[−3.1,+2.8]
Appendix
Table 19 : Pass-rate offsets (percentage points) on instances from Doubao Seed and other source models. An offset is a response’s pass indicator minus the mean pass rate of the other eight LLMs on the same instance. The difference is Doubao-sourced minus other-sourced; 95% confidence intervals resample instances or benchmarks.
Analysis
Subset
Why this subset
Model pass rates overall and by frame and setting (Table 2 )
107 benchmarks, 4,065 instances
Instances with valid responses from all nine LLMs, so that the LLMs are compared on the same instances; 60 of the 4,125 constructed instances lack a valid response from at least one LLM.
Uncertainty in model rankings (Table 18 )
107 benchmarks, 4,065 instances
The instances of the main comparison; benchmarks are resampled with replacement.
Behavior groups (Figures 4 a and 7 ; Tables 24 , 25 , and 17 )
102 benchmarks, 3,966 instances
Excludes the five AND benchmarks, which test two behaviors and cannot be assigned to one group.
Behavior families (Figures 5 and 4 c; Table 9 )
87 benchmarks, 3,444 instances
Single-behavior benchmarks built from catalog specifications; benchmarks built from query-specific specifications belong to no catalog family.
First-action comparisons (Figure 4 b; Table 26 )
Varies by action (259 instances for TodoWrite )
Each comparison uses the instances on which some LLMs choose the action and others do not.
Judge self-preference and judge removal (Tables 15 and 16 )
102 benchmarks, 3,945 instances
Single-behavior instances with valid scores from all three judges for all nine LLMs.
Appendix
Table 20 : Benchmarks, instances, and queries used in each analysis.
Query type
Expected outcome
Queries
Actual outcome
Built
Rejected
Unmet
Predefined
Build
98
94
0
4
Missing parameter
Reject
5
0
5
0
Custom
Build
9
8
0
1
Adversarial
Reject
27
0
27
0
Total
–
139
102
32
5
Appendix
Table 21 : Expected and actual outcomes by query type. “Unmet” denotes build requests that did not produce all requested benchmarks.
Query
Failure point
What prevented completion
Respect for rejection during unattended continuation
Specification checks
The query requires a user rejection followed by unattended continuation, but the specification excludes all sessions with user-written input.
Hard-coded home-directory paths
Anchor Synthesis Loop
The synthesized specifications could not reliably distinguish inappropriate hard-coded paths from acceptable ones.
Error-information use in unattended sessions
Anchor Synthesis Loop
The synthesized specifications could not distinguish ignoring available information from legitimate re-checking of that information.
Error-guided correction OR Retry adaptation
Specification matching
Error-guided correction produced 50 instances, but specification matching failed for Retry adaptation, leaving the second benchmark unbuilt.
Regression recognition OR Conflict-aware editing
Benchmark construction
Regression recognition produced six instances, but no benchmark was built for Conflict-aware editing.
Appendix
Table 22 : Why five build requests were unmet. OR queries require successful construction of both benchmarks.
Measure
Mean
Median
Successful queries ( n=102 )
Anchor session scans
128,297
130,152
Candidates reviewed by the Fast Model
706
292
LLM calls
552
395
Input and output tokens (millions)
25.14
16.44
Rejected queries
Appendix
Table 23 : Mean and median benchmark-construction workload per query.
Table 24 : Behaviors in the four groups used for analysis. Predefined behavior names follow Table 9 ; query-specific names summarize their passing requirements.
Specifications included
Delivery recovery group
Valid call
Handle failure
Honest claim
Check first
Predefined + query-specific
Handle failure
67.9
33.5
28.9
8.1
Predefined + query-specific
Check first
67.9
36.7
28.9
7.7
Predefined only
Handle failure
69.1
33.3
28.9
9.4
Predefined only
Check first
69.1
36.8
28.9
8.7
Appendix
Table 25 : Pass rates (%) after changing group membership or using only predefined specifications.
First action
Instances
First responses
Other responses
First pass rate
Other pass rate
Δ
TodoWrite
259
457
1,874
4.6
30.8
−26.2
EnterPlanMode
61
65
484
9.2
28.1
−18.9
Skill
395
419
3,136
13.4
29.7
−16.3
Text only
1,450
4,662
8,388
11.9
24.7
−12.8
AskUserQuestion
128
229
923
46.3
25.1
+21.2
Agent (subagent)
170
235
1,295
55.3
31.5
+23.8
Appendix
Table 26 : First-action comparisons. First/other denote responses choosing the listed action or any alternative on the same instances. Rates are percentages; Δ is first minus other in percentage points. Text-only responses contain no tool call. Figure 4 b shows the first actions with ∣Δ∣>10 .
Table 27 : Recorded context and evaluation cut point.
Table 28 : Anchor match, the source LLM’s next response, and Fast Model confirmation. These records are used only during construction and are not part of the evaluation input.
Table 29 : Responses of Kimi-K3 and Claude Opus 4.8 at the cut point. Each response consists of a single tool call without text.
Table 30 : Scoring rubric for Error-guided correction .
Evaluated LLM
Judge
Score
Rationale
Kimi-K3
GPT-5.6-Sol
4
Correctly creates and uses a virtual environment and reruns the script, but also installs akshare , which this script does not need; slightly redundant.
Gemini-3.5-Flash
4
Creates a virtual environment and uses source activation, which is correct but can sometimes be fragile in non-interactive shells.
Claude Opus 4.8
5
Builds a venv and installs there, exactly the error’s primary recommended path.
Claude Opus 4.8
GPT-5.6-Sol
2
Only probes the modules and the skill directory, and does not carry out the virtual-environment or override-flag fix the error explicitly gave.
Gemini-3.5-Flash
2
Performs redundant diagnostics checking for packages that the environment already explicitly reported as missing.
Claude Opus 4.8
2
Probes whether packages are already importable instead of applying the venv fix.
Appendix
Table 31 : Recorded judge scores and rationales for the two responses.
Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fulfilled. CAR-bench exposes this reliability gap in the domain of in-car assistants: an LLM-simulated user issues incomplete or ambiguous requests, requiring the agent to resolve uncertainty through multi-turn dialogue and tool use while strictly adhering to domain policies. Even frontier models show a substantial gap between what they can solve at least once (Pass@3) and what they solve consistently across trials (Pass^k). We bridge this gap with TRACE (TRAjectory-Contrastive Evolution), which iteratively improves a skill-based agent's behavioral knowledge without modifying model weights. This knowledge is organized as a Skill Bank of modular, retrievable skills, each encoding a self-contained set of tool-use rules and behavioral guidelines. TRACE evolves this bank through an agentic self-evolution loop: after each evaluation round, it groups trajectories by the skills invoked and refines each skill by contrasting successful and failed behaviors. The updated bank then guides subsequent rounds, while during deployment the Actor performs state-conditioned skill orchestration at every turn. On GPT-5.5, TRACE improves consistency (Pass^3) by 34.6 points, from 59.9% to 94.5%, while shrinking the gap between potential and reliable performance to just 4.0 points. On the official hidden set, TRACE achieved first place using GPT-5.6-Sol, attaining a Pass^3 score of 70%-a 40% relative improvement over the baseline. These results show that TRACE converts high model potential into stable, consistent performance gain. Project homepage: https://darwin-agent.github.io/Car-bench-TRACE.
Wenhao Wu, Menghao Zhang, Xin Wang +3
Xiaomi Inc. · Nanjing University · Tsinghua University
Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, leaving reflection quality unknown, and it relies on agents' own episode runs, offering no mechanism to target specific failure patterns. We present \textbf{BenchTrace}, a benchmark for evaluating self-evolution ability in LLM agents. BenchTrace is built on a snapshot-reflection dataset of 1,821 annotated episodes spanning six diverse tasks, and comprises a \textbf{Reflection Evaluation} that probes failure identification through targeted QA tasks, and an \textbf{Evolution Evaluation} that tests whether past failure experience translates into avoidance behavior in a controlled self-evolution simulation. Building on BenchTrace, we propose \textbf{failure avoidance rate (FAR)}, a new evaluation metric measuring the fraction of test cases in which the agent successfully avoids the target failure instance. Experiments with Qwen3-32B and GPT-4.1 reveal that both models fall below a 30% end-to-end pass rate on reflection evaluation, with diagnosis as the primary bottleneck. Evolution evaluation shows that self-evolution methods generally improve FAR over the non-evolving baseline, but agents forget early lessons as noise episodes accumulate, and agents fail to generalize their reflections beyond the specific context, causing negative transfer across task contexts. Our correlation analysis further reveals that only a fully correct reflection is strongly associated with higher FAR. BenchTrace exposes concrete limits of current self-evolution approaches and provides a controlled, model-agnostic framework for targeted evaluation.
Jiahao Huang, Fei Cheng, Junfeng Jiang +2
University of Tokyo · Kyoto University · National Institute of Informatics
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.