ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Organizations: Uniphore · University of Illinois Urbana-Champaign
Abstract
Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as -bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across -bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
Figures & tables
| Domain | H | U | I | Total |
|---|---|---|---|---|
| Train | ||||
| Banking | 260 | 208 | 203 | 671 |
| Calendar assistant | 199 | 199 | 188 | 586 |
| Home services | 186 | 209 | 206 | 601 |
| Online shopping | 215 | 286 | 211 | 712 |
| Restaurant booking | 204 | 204 | 201 | 609 |
| Dialogue statistics | Coherence | Diversity | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | Tools per dialogue | Turns per dialogue | User turns per dialogue | Calls per user turn | Semantic similarity | Entailment rate | Entropy | Distinct-3 | Dialogue Vendi | Tool-trajectory Vendi |
| APIGen-MT-5k | 15.1 | 18.5 | 4.8 | 0.91 | 0.504 | 0.071 | 8.62 | 0.200 | 18.8 | 12.1 |
| ToolRACERBench | 11.8 | 12.6 | 3.5 | 0.82 | 0.559 | 0.042 | 9.29 | 0.348 | 23.8 | 18.8 |
| BFCL v3 | -Bench | ||||||||
| Single Turn (%) | Hallucination (%) | Pass@1 (%) | |||||||
| Model / Training setting | Non-live Overall | Multiple | Parallel | Parallel Multiple | Irrelevance | Relevance | Airline | Retail | Macro |
| Open Source Models | |||||||||
| Qwen3-4B-Instruct | 89.83 | 91.00 | 90.00 | 88.50 | 85.83 | 87.50 | 24.0 | 40.4 | 32.20 |
| ToolACE-MT | 84.94 | – | – | – | 72.83 | 77.78 | 16.0 † | 25.2 † | 20.60 † |
| xLAM-2-3B-FC-R | 82.94 | – | – | – | 57.94 | 94.44 | 32.0 † | 44.4 † | 38.20 † |
| Happy (%) | Unhappy (%) | Impossible (%) | ||||||||||
| Model / Training setting | AR | Tool F1 | Params | AnsSim | AR | Tool F1 | Params | AnsSim | AR | Tool F1 | Params | AnsSim |
| Open Source Models | ||||||||||||
| Qwen3-4B-Instruct | 75.9 | 72.8 | 43.2 | 51.9 | 77.3 | 72.8 | 41.1 | 54.1 | 68.6 | 77.8 | 45.0 | 44.5 |
| xLAM-2-3B-FC-R | 75.7 | 61.2 | 35.6 | 56.4 | 75.4 | 54.2 | 27.6 | 56.1 | 66.0 | 62.1 | 31.4 | 39.6 |
| Fine-tuned Qwen3-4B-Instruct – Main Results | ||||||||||||
| APIGen-MT | 79.3 | 62.1 | 39.0 | 67.5 | 80.7 | 61.3 | 35.3 | 67.8 | 79.8 | 66.9 | 40.6 | 63.9 |
| Accuracy (%) | ||
| Model / Training setting | EA | PA |
| Open-source models | ||
| ToolACE-8B | 6.7 | 27.7 |
| xLAM2-3b-fc | 0.0 | 14.7 |
| Base models and fine-tuned variants | ||
| Qwen3-4B-Instruct-2507 | 8.4 | 29.2 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Success Rate (%) | ||||
|---|---|---|---|---|
| Metric | Happy | Unhappy | Impossible | All |
| Multi-Agent Simulation (w/o refinement) | 32.2 | 23.9 | 23.6 | 26.3 |
| Refinement Recovery Rate | 72.6 | 62.3 | 75.0 † | 65.4 |