Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as τ2-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across τ2-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
Figures & tables
Figure 1: Examples of ToolRACERBench conversation trajectories: (a) an Unhappy path and (b) an Impossible path.
Figure 2: Overview of ToolRACER for robust data generation. Scenario specifications condition interactions among user, assistant, and tool agents, with the assistant choosing to act through tool calls or respond in natural language. Four judges assess syntax errors, faithfulness, role confusion, and task success. Validated conversations are retained in the dataset, while failure reasons serve as corrective hints to repair and re-simulate unsuccessful conversations. Cases that remain invalid after the maximum number of attempts are discarded.
Domain
H
U
I
Total
Train
Banking
260
208
203
671
Calendar assistant
199
199
188
586
Home services
186
209
206
601
Online shopping
215
286
211
712
Restaurant booking
204
204
201
609
Table 1: ToolRACERBench split by domain and path type. H/U/I denote Happy / Unhappy / Impossible .
Dialogue statistics
Coherence
Diversity
Dataset
Tools per dialogue
Turns per dialogue
User turns per dialogue
Calls per user turn
Semantic similarity ↑
Entailment rate ↑
Entropy ↑
Distinct-3 ↑
Dialogue Vendi ↑
Tool-trajectory Vendi ↑
APIGen-MT-5k
15.1
18.5
4.8
0.91
0.504
0.071
8.62
0.200
18.8
12.1
ToolRACERBench
11.8
12.6
3.5
0.82
0.559
0.042
9.29
0.348
23.8
18.8
Table 2: Dataset statistics and quality metrics for APIGen-MT-5k and ToolRACERBench. Bold indicates the higher value in each column; higher structural counts do not necessarily indicate better quality.
BFCL v3
τ2 -Bench
Single Turn (%) ↑
Hallucination (%) ↑
Pass@1 (%) ↑
Model / Training setting
Non-live Overall
Multiple
Parallel
Parallel Multiple
Irrelevance
Relevance
Airline
Retail
Macro
Open Source Models
Qwen3-4B-Instruct
89.83
91.00
90.00
88.50
85.83
87.50
24.0
40.4
32.20
ToolACE-MT
84.94
–
–
–
72.83
77.78
16.0 †
25.2 †
20.60 †
xLAM-2-3B-FC-R
82.94
–
–
–
57.94
94.44
32.0 †
44.4 †
38.20 †
Table 3: Performance on BFCL v3 and τ2 -Bench across training settings, with Happy (H), Unhappy (U), and Impossible (I) ablations. All scores are percentages; higher is better. Bold denotes the best available score in each column, including ties. Blue shading marks ToolRACERBench-based main results. “–” denotes unavailable results. For BFCL, non-live overall accuracy is reported independently of the displayed categories. For τ2 -Bench, Macro is the unweighted mean of Airline and Retail pass@1; † reported results may use different evaluation protocols.
Happy (%) ↑
Unhappy (%) ↑
Impossible (%) ↑
Model / Training setting
AR
Tool F1
Params
AnsSim
AR
Tool F1
Params
AnsSim
AR
Tool F1
Params
AnsSim
Open Source Models
Qwen3-4B-Instruct
75.9
72.8
43.2
51.9
77.3
72.8
41.1
54.1
68.6
77.8
45.0
44.5
xLAM-2-3B-FC-R
75.7
61.2
35.6
56.4
75.4
54.2
27.6
56.1
66.0
62.1
31.4
39.6
Fine-tuned Qwen3-4B-Instruct – Main Results
APIGen-MT
79.3
62.1
39.0
67.5
80.7
61.3
35.3
67.8
79.8
66.9
40.6
63.9
Table 4: ToolRACERBench-Test evaluation across Happy (H), Unhappy (U), and Impossible (I) paths. All scores are expressed as percentages; higher is better. Bold denotes the best score in each column, including ties. Blue shading marks ToolRACERBench-based main results. AR: Action Recall; Params: Full Parameter Match; AnsSim: Agent Answer Similarity.
Accuracy (%) ↑
Model / Training setting
EA
PA
Open-source models
ToolACE-8B
6.7
27.7
xLAM2-3b-fc
0.0
14.7
Base models and fine-tuned variants
Qwen3-4B-Instruct-2507
8.4
29.2
Table 5: ACEBench agent performance (%), averaged across multi-step and multi-turn tasks. EA: end-to-end accuracy; PA: process accuracy. Blue shading denotes ToolRACERBench-based fine-tuning; bold marks the best score per model size.
Figure 3: Prompt used during ToolRACER training and data generation, for the Tool agent initialized in the Multi Agent Simulation.
Figure 4: Prompt used during ToolRACER training and data generation, for the User agent initialized in the Multi Agent Simulation. Because this prompt is too long to fit on one page, we break it up into two. The rest of the prompt is in Figure 5 .
Figure 5: Prompt used during ToolRACER training and data generation, for the User agent initialized in the Multi Agent Simulation. Because this prompt is too long to fit on one page, we break it up into two. The prefix of the prompt is in Figure 4 .
Figure 6: Prompt used during ToolRACER training and data generation, for the Faithfulness Judge initialized in the Multi Stage Evaluation.
Figure 7: Prompt used during ToolRACER training and data generation, for the Role Confusion Judge initialized in the Multi Stage Evaluation.
Figure 8: Happy metadata.
Figure 9: Unhappy metadata
Figure 10: Impossible metadata.
Figure 11: Tools from the domain of home services
Figure 12: Example of one persona used in ToolRACERBench.
Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain limited in task complexity, realism, and domain diversity, and often fail to capture interactions that span multiple domains, limiting their ability to evaluate agents in realistic multi-step settings that require sustained reasoning and coordination. To address these limitations, we introduce T1-Bench, a high-fidelity, comprehensive benchmark for evaluating agentic systems in realistic customer-facing, multi-domain environments, featuring interleaved scenarios that require structured reasoning across multi-turn user-assistant interactions and substantially increasing both compositional complexity and evaluative rigor across 25 domains of varying difficulty. We evaluate T1-Bench using 12 proprietary and open-weight models, providing a reproducible and standardized framework for assessing agent behavior, tool utilization, and conversational quality in complex, multi-step environments. We further complement automatic evaluation with human judgments to strengthen the assessment of qualitative performance. Overall, T1-Bench substantially advances prior benchmarks by increasing task complexity, interaction depth, and domain coverage in simulated multi-domain environments. To facilitate future research on agentic systems, we will publicly release data and evaluation code as open source.
Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user typos propagate into hallucinated tool names, a misconfigured request timeout can stall an agent indefinitely, and duplicate tool names across servers can freeze an SDK. We study these failures as a sim-to-real gap in the tool-use partially observable Markov decision process (POMDP), where deployment noise enters through the observation, action space, reward-relevant metadata, or transition dynamics. We introduce RobustBench-TC, a benchmark with 22 perturbation types organized by these four POMDP components, each grounded in a verified GitHub issue or documented tool-calling failure. Across 21 models from 1.5B to 32B parameters (including the closed-source o4-mini), the robustness profile is sharply uneven: observation perturbations reduce accuracy by less than 5%, while reward-relevant and transition perturbations reduce accuracy by roughly 40% and 30%, respectively; scale alone does not close these gaps. We then propose ToolRL-DR, a domain-randomization reinforcement learning (RL) recipe that trains a tool-use agent on perturbation-augmented trajectories spanning the three statically encodable POMDP components. On a 3B backbone, ToolRL-DR-Full retains roughly three-quarters of clean accuracy and reaches an aggregate perturbed accuracy comparable to open-source 14B function-calling baselines while substantially narrowing the gap to o4-mini. It closes approximately 27% of the Transition gap despite never seeing transition perturbations in training, suggesting that RL on adversarial static tool-use inputs induces a more persistent retry policy that transfers to unseen runtime failures. The dataset, code and benchmark leaderboard are publicly available.
Xiaolin Zhou, Aojie Yuan, Zheng Luo +12
Arizona State University · University of Southern California · University of Pennsylvania +2
Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and tools. While existing evaluation methods assess either user satisfaction or agents' tool-calling capabilities, they fail to capture critical errors in multi-turn tool-augmented dialogues-such as when agents misinterpret tool results yet appear satisfactory to users. We introduce TRACE, a benchmark of systematically synthesized tool-augmented conversations covering diverse error cases. Evaluation with state-of-the-art conversation evaluation frameworks reveals that all approaches remain far from ideal performance, demonstrating the fundamental difficulty of this benchmark.