Turnslide: Scalable Multi-Turn Data Synthesis by Walking a Finite-State Machine
Organizations: Imperial College London · distil labs · Agon · ELLIS Institute Tübingen
Abstract
Small language models are inexpensive to serve and can run on private infrastructure, but base models are often not good enough at multi-turn tool calling, and fine-tuning them needs per-API data that rarely exists. Existing synthesis methods are too expensive for high-scale fine-tuning, as they often require mock operational environments for different domains and multiple LLM calls per generated conversation turn. We introduce a fully automated, lightweight synthesis framework that models each API as a finite-state machine, representing the system as abstract states that determine when each tool may be called, producing state-valid sequences of tools; sequences are translated into complete examples with a single LLM call. Rather than optimize diversity, we set a target distribution over the number of turns, the tool sequence and task complexity. We measure data quality by fine-tuning SLMs on generated trajectories, showing that our FSM-based generation significantly improves downstream accuracy over an unmutated baseline and, against existing works, reaches 70.7% full accuracy over 63.4% and 53.7% with 3.6-6.6 fewer tokens.
Figures & tables
| Method | Summary | Trading | Travel | VehicleControl | FileSystem 1 | Tick.&Trav. 2 | ||||||
| Per Turn | Full | Per Turn | Full | Per Turn | Full | Per Turn | Full | Per Turn | Full | |||
| Base Student | — | — | 42.7 | 26.3 | 48.0 | 21.5 | 38.5 | 5.6 | 49.4 | 18.9 | 34.3 | 12.5 |
| Baseline | — | — | 76.9 | 43.9 | 79.6 | 47.8 | 69.4 | 18.8 | 80.7 | 67.3 | 78.1 | 38.5 |
| Turns Mutator | 2.89 2.87 | 0.321 | 76.0 | 44.9 | 76.5 | 42.3 | 64.3 | 19.1 | 88.9 | 74.2 | 76.2 | 40.4 |
| Markovian | 7.88 3.19 | 0.03 | 76.8 | 41.5 | 83.1 | 60.6 | 70.3 | 19.7 | 94.9 | 84.5 | 75.9 | 36.6 |
| FSM | 10.1 2.62 | 0.01 | 81.5 | 54.2 | 84.2 | 61.5 | 72.6 | 22.5 | 94.5 | 80.2 | 78.2 | 47.2 |
| Method | Calls/ Trajectory | Tokens/ Trajectory | Tokens/ Turn | Examples/ 20M Tokens | Fine-Tune Full Acc.(%) | Fine-Tune Per Turn Acc. (%) | LLM Eval (/5) |
| ToolFlow | 34.3 | 99,846 | 22,327 | 200 | 63.4 | 85.7 | 3.03 |
| APIGen-MT | 107.2 | 184,304 | 35,943 | 109 | 53.7 | 80.1 | 3.04 |
| Ours | 1.0 | 27,724 | 5,545 | 721 | 70.7 | 87.2 | 3.34 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Model | Trading | Travel | VehicleControl | FileSystem 1 | Tick.&Trav. 2 | |||||
| Per Turn | Full | Per Turn | Full | Per Turn | Full | Per Turn | Full | Per Turn | Full | ||
| Base Student | Qwen | 79.9 | 52.6 | 62.3 | 30.0 | 53.9 | 9.9 | 68.7 | 25.3 | 55.6 | 23.1 |
| Llama | 5.5 | 0.0 | 33.6 | 13.0 | 23.1 | 1.3 | 30.0 | 12.5 | 13.0 | 1.9 | |
| Baseline | Qwen | 76.0 | 39.0 | 74.8 | 38.5 | 64.9 | 10.7 | 82.3 | 65.5 | 72.6 | 30.7 |
| Llama | 77.7 | 48.7 | 84.4 | 57.1 | 73.8 | 26.9 | 79.1 | 69.0 | 83.5 | 46.2 | |
| Turns Mutator | Qwen | 73.8 | 38.3 | 76.5 | 42.3 | 64.1 | 13.1 | 89.2 | 72.4 | 69.6 | 30.8 |