TokenRouter: Efficient Serving System for Token-Level LLM Routing
Organizations: Tsinghua University · Carnegie Mellon University
Abstract
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at https://github.com/thu-nics/TokenRouter.
Figures & tables
| Algorithm | #Models | Model | Model |
| CITER [ 12 ] | 2 | low-confidence token | after one token |
| R2R [ 5 ] | 2 | predicted-divergent token | after one token |
| R-Stitch [ 27 ] | 2 | high-entropy token | low-entropy token |
| Co-LLM [ 10 ] | 2 | high-deferral token | after one token |
| ME [ 33 ] | 2 | after every token, according to ensemble weights | |
| Algorithm | Implementation | Throughput/(token/s) | TTFT/s | Latency/s |
| R2R | LLM-only | 145.30 | 0.13 | 411.98 |
| Official Code | 89.62 | 0.11 | 751.15 | |
| TokenRouter | 244.56 | 0.11 | 270.19 | |
| CITER | LLM-only | 123.72 | 0.083 | 2.68 |
| Official Code | 17.16 | 0.036 | 7.64 | |
| TokenRouter | 149.31 | 0.036 | 0.48 |
| System | Model Pair | Benchmark |
| R2R | DeepSeek-R1-Distill-Qwen-1.5B / 32B | AIME [ 35 ] |
| CITER | Qwen2-1.5B / 72B | CommonsenseQA [ 37 ] |
| Co-LLM | LLaMA2-7B (tuned) / 70B | GSM8K [ 38 ] |
| R-Stitch | L1-1.5B-Short / QwQ-32B | AIME [ 35 ] |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Algorithm | Implementation | Throughput/(token/s) | TTFT/s | Latency/s |
| R2R | LLM-only | 40.17 | 0.074 | 388.12 |
| Official Code | 54.77 | 0.12 | 330.69 | |
| TokenRouter | 105.81 | 0.11 | 144.10 | |
| CITER | LLM-only | 34.40 | 0.055 | 2.42 |
| Official Code | 14.26 | 0.031 | 2.41 | |
| TokenRouter | 50.63 | 0.031 | 0.36 |
| Method | |||
| TokenRouter | 372.48 | 210.61 | 73.02 |
| - Delayed batching | 296.86 | 161.26 | 73.02 |
| - Async execution | 230.79 | 143.88 | 73.02 |
| - Router CUDA graph | 228.10 | 142.33 | 72.62 |
| - LLM CUDA graph | 164.08 | 95.55 | 51.22 |
| - SLM CUDA graph ( R2R) | 132.78 | 75.54 | 34.20 |
| Total | #SLM | #LLM | Overlap | ||
| 2 | 1 | 1 | 50.37 | 160.74 | |
| 1 | 2 | 70.83 | 197.40 | ||
| 2 | 2 | 71.41 | 198.13 | ||
| 8 | 4 | 4 | 93.23 | 284.62 | |
| 1 | 8 | 109.67 | 296.34 | ||
| 2 | 8 | 106.02 | 289.08 |
| Setup | Routing Algorithm | ||||
| Deployment | R2R | R-Stitch | Co-LLM | CITER | |
| 1 | Overlapping | 73.27 | 43.06 | 47.46 | 102.29 |
| Non-overlapping | 74.90 | 47.53 | 48.25 | 108.65 | |
| 4 | Overlapping | 205.05 | 108.42 | 165.96 | 281.27 |
| Non-overlapping | 228.94 | 116.14 | 170.20 | 308.36 | |
| Setup | Routing Algorithm | |||
| Deployment | R-Stitch | R2R | CITER | |
| 1 | Single-node | 47.53 | 74.90 | 108.65 |
| Multi-node | 46.19 | 69.92 | 93.99 | |
| 4 | Single-node | 116.14 | 228.94 | 308.36 |
| Multi-node | 104.20 | 202.43 | 277.56 | |
| Operation | Latency (ms) | Ratio |
| Prefix matching | 7.78 | 20.94% |
| Locking cache nodes | 6.49 | 17.47% |
| SLM inference | 1.56 | 4.20% |
| Releasing locks | 5.13 | 13.81% |
| Update radix cache | 12.12 | 32.62% |
| Others | 4.08 | 10.98% |
| Notation | Definition |
| Number of cooperating models. | |
| Number of concurrent requests in the system. | |
| Latency of one decoding step on model . | |
| Probability that a request is routed from model to model after a decoding step. | |
| Number of output tokens committed when a request is routed from model to model . | |
| Delayed-batching threshold of model . |