Small language models (SLMs) are attractive as local agent controllers because they reduce remote inference, latency, and deployment footprint, yet structured tool errors can cause an agent step to fail. Existing routers typically select a model once per query. However, agents expose sequential decision points whose difficulty dynamically changes based on intermediate observations. We propose STEPGATE, an uncertainty-aware handoff framework that scores each local SLM action and selectively escalates challenging steps to a stronger model. On a 52-task held-out single-step BFCL-derived test split, the Qwen2.5-1.5B/7B pair attains 82.7% task success with 30.8% escalation, versus 67.3% local-only and 75.4% random escalation (which uses 33.8% escalation). In a separate multi-turn evaluation, STEPGATE achieves 69.0% trajectory success and 84.0% action success using only 30.0% cloud actions, compared with 48.0%/70.5% local-only, 60.0%/78.2% random escalation, and 57.0%/77.1% query-level routing (strong-only achieves 82.0% trajectory success at 100% cloud actions). These results suggest that step-level escalation recovers a large share of the performance gap to the stronger Qwen2.5-7B backend at a matched cloud-action rate while transmitting fewer tokens remotely. However, our evaluation is limited to one model family, a single stronger backend, and scripted tasks. Furthermore, the test sets are small, multi-turn comparisons rely on paired intervals and statistical tests, and our risk tiers serve as research annotations rather than formal safety guarantees.
Figures & tables
Figure 1: StepGate routes each agent step . The local SLM first proposes an action. A lightweight gate estimates failure probability from action-structured uncertainty and risk. Only uncertain/high-risk steps disclose the required context to a stronger model.
Method
Small → large
Step
Structured UQ
Tool risk
Agent metric
Guarantee
FrugalGPT ( Chen et al., 2024 )
✓
–
–
–
–
–
RouteLLM ( Ong et al., 2025 )
✓
–
–
–
–
–
Hybrid LLM ( Ding et al., 2024 )
✓
–
–
–
–
–
AutoMix ( Aggarwal et al., 2024 )
✓
–
partial
–
–
–
UALA ( Han et al., 2024 )
–
partial
–
–
–
–
RACER ( Hao et al., 2026 )
✓
–
–
–
–
✓
Table 1: Positioning against representative recent work. “Step” means routing can change after each tool observation. “Structured UQ” uses tool/argument/state signals. “Risk” accounts for action consequence. “Guarantee” denotes a finite-sample/distribution-free risk-control guarantee.
Figure 2: Training and deployment. Failure labels supervise only the gate. Thresholds are selected on validation data to satisfy a cloud-call/risk budget, enabling deployment-time trade-offs without retraining either language model.
System
Task succ. ↑
Cloud % ↓
Wrong risky ↓
Latency ↓
Remote tok. ↓
SLM only
67.3
0
15.4
0.702
0
Strong only
92.3
100
3.8
2.255
475.0
Random @ budget
75.4
33.8
10.8
0.823
158.5
Query-level router
73.1
42.3
15.4
1.052
266.2
Self-consistency
82.7
38.5
7.7
0.829
175.2
StepGate
82.7
30.8
3.8
0.777
134.5
Table 2: Main results: qwen2.5-1.5b (local) / qwen2.5-7b 4-bit (strong), real 100-task BFCL subset, 52 test tasks. Cloud rate budget-matched at 30% target (threshold/model chosen on calibration split only, applied once to test). “Wrong risky” = % of executed test actions that were both incorrect and risk-tier ≥ R2 (our taxonomy has no R3 action type in this task set). Realized cloud rates differ across routers and rate-matched variants are in Appendix A.1 .
Method
Traj. succ. ↑
95% CI ‡
Action succ. ↑
Cloud actions ↓
Mixed success ↑
Local-only
48.0%
[38.5, 57.7]
70.5%
0.0%
–
Random escalation
60.0%
[50.2, 69.1]
78.2%
30.0%
34.0%
Query-level router
57.0%
[47.2, 66.3]
77.1%
30.0%
0.0%
Self-consistency
66.0%
[56.3, 74.5]
82.6%
38.0%
48.0%
StepGate
69.0%
[59.4, 77.2]
84.0%
30.0%
55.0%
Strong-only
82.0%
[73.3, 88.3]
91.0%
100.0%
–
Table 3: Multi-turn results for Qwen2.5-1.5B local and Qwen2.5-7B strong models, evaluated on the 100-task multi-turn split described in Appendix A.6 . “Mixed success” is the share of trajectories that combined at least one local and one cloud action and still succeeded. ‡ Unpaired 95% Wilson interval for each method’s trajectory success ( n=100 ; random escalation is the five-seed mean, so its interval is approximate). Paired differences are given in the text.
Figure 3: Measured end-to-end task success versus realized strong-model step escalation rate for Qwen2.5-0.5B, 1.5B, and 3B local models. Error bars are the task-level 95% confidence intervals from the executed routing sweep. The dashed oracle is non-deployable. The plot exposes an important scale effect: routing offers substantially more headroom for weaker local models, whereas the 3B local baseline is already close to the strong-model ceiling.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Realized esc.
Task success
95% Wilson CI
Wrong R2+
Local-only
0.0%
67.3%
[53.8, 78.5]
15.4%
Random escalation, as reported †
33.8%
75.4%
[62.2, 85.1]
10.8%
Random escalation, matched †
30.8%
74.7%
[61.5, 84.5]
11.2%
StepGate (full)
30.8%
82.7%
[70.3, 90.6]
3.8%
StepGate no-risk, as reported
34.6%
86.5%
[74.7, 93.3]
3.8%
StepGate no-risk, matched
30.8%
82.7%
[70.3, 90.6]
5.8%
Appendix
Table 4: Single-step test (52 tasks), Qwen2.5-1.5B/7B. † Five-seed mean; interval approximate. “As reported” rows use the thresholds of the original experiment. “matched” rows are re-tuned on validation data to 30.8% realized escalation.
Gate features
AUROC ↑
ECE ↓
Task succ. ↑
Cloud %
All ( StepGate )
0.829
0.212
82.7
30.8
− consistency
0.832
0.231
82.7
28.8
− schema/arguments
0.829
0.241
82.7
30.8
− risk
0.886
0.119
86.5
34.6
− risk, matched budget
0.886
0.119
82.7
30.8
Token uncertainty only
0.565
0.078
76.9
28.8
Appendix
Table 5: Ablation of the calibrated failure predictor (logistic regression), qwen2.5-1.5b, real BFCL subset (n=52 test tasks). Cloud budget fixed at 30% (threshold chosen on calibration split only). The risk gate escalates 2 more of 52 steps than the full gate. The last-but-one row re-tunes its threshold to the same 30.8% budget (AUROC/ECE are unchanged because the scoring model is the same).
Figure 4: Adaptive uncertainty acquisition. Expensive self-consistency is computed only for ambiguous actions. This appendix figure makes explicit which uncertainty signals incur additional local inference. Note that this diagram covers only sample-based consistency and a single bounded revision.
read-only/state-changing/destructive/external side effect
No
† Evaluated exclusively within multi-turn sequential trajectories.
Appendix
Table 6: Feature taxonomy and reporting overview. “Extra decode” identifies signals that incur additional local generation and therefore must be accounted for in latency/energy.
delete_event(recurring=True) ; permanently_delete_file(...) ; bulk_delete_emails(filter=...) ; cancel_meeting(notify_all=True) for an externally organized event
Appendix
Table 7: Worked examples of the risk taxonomy used to annotate tool calls. Tier assignment follows reversibility and blast radius, not tool name alone: two calls to the same underlying tool can receive different tiers depending on arguments (e.g., a narrowly scoped read vs. a bulk export).
Local / strong pair
Params
Task succ.
Cloud %
Risky errors
Latency
qwen2.5-0.5b / qwen2.5-7b
0.5B
73.1
26.9
3.8
0.351
qwen2.5-1.5b / qwen2.5-7b
1.5B
82.7
30.8
3.8
0.777
qwen2.5-3b / qwen2.5-7b
3B
90.4
26.9
3.8
1.153
Appendix
Table 8: Per-model results at 30% fixed cloud budget, real BFCL subset. Three local/strong pairs were evaluated (a single strong model and three SLM sizes).
Figure 5: Selective risk under entropy-ranked local coverage. The error rate among locally executed steps rises with coverage, most sharply for the 0.5B model and least for the 3B model. This plot evaluates ranking behavior; it is not the learned-gate routing curve.
Figure 6: Reliability diagrams for the learned logistic failure gate. From left to right, the panels show Qwen2.5-0.5B, Qwen2.5-1.5B, and Qwen2.5-3B. The dashed diagonal denotes perfect calibration. The finite calibration/test sample sizes make individual bins noisy, so ECE and Brier scores are reported alongside these diagrams rather than inferred visually.
Figure 7: Measured task success versus mean strong-model tokens per task (input+output). Each point is an executed routing configuration. The dashed connections are visual guides, not fitted performance models. The figure shows that token efficiency and task success must be considered jointly and that the preferred operating point depends on local model size.
Figure 8: Measured per-call latency, peak GPU memory, and NVML power-draw integration for the local SLMs and the shared 7B 4-bit strong model. These measurements are specific to the hardware/runtime in use. They are not hardware-independent efficiency claims.
Figure 9: Failure decomposition and escalation recoverability for the Qwen2.5-0.5B local model (left), Qwen2.5-1.5B local model (middle), and Qwen2.5-3B local model (right). The categories include irrelevance/abstention errors, malformed outputs, parallel-call count or matching errors, missing optional or required arguments, and value errors involving strings, lists/tuples, or other values. Red bars show observed local failure counts on the test and calibration sets, while the green line shows the recovery rate when the same failing step is escalated to the strong model. Recovery varies sharply by failure type, underscoring why local failure probability alone is not equivalent to the expected escalation benefit. Sparse categories should be interpreted cautiously.
SLM
Local in tok.
Local out tok.
Strong in tok.
Strong out tok.
Local lat. p50/p95 (s)
Strong lat. p50/p95 (s)
Local peak mem (MiB)
Local energy p50 (J)
qwen2.5-0.5b
427.9
41.1
427.9
47.1
0.301/0.576
1.528/3.385
990
24.31
qwen2.5-1.5b
427.9
47.9
427.9
47.1
0.702/1.638
1.528/3.385
3003
80.66
qwen2.5-3b
427.9
43.9
427.9
47.1
1.152/2.248
1.528/3.385
6058
143.86
Appendix
Table 9: Measured efficiency statistics from the executed runs. Strong-model latency is repeated by row because the same Qwen2.5-7B 4-bit model is used.
Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM cascades usually route whole queries before execution, but task difficulty shifts mid-trajectory - after flaky tool calls, truncated observations, or compounding local errors - making pre-execution routing brittle. We introduce \textbf{R2V-Agent}, a risk-calibrated SLM-LLM routing framework for interactive agents. R2V combines four components: a distilled small language model (SLM) policy, a stronger teacher LLM, a lightweight process verifier that scores candidate actions at each step, and a calibrated step-level router. The router is our central contribution: after the SLM is trained, it estimates residual failure risk at each step and escalates only when teacher intervention is warranted. To make the routing problem well-defined, we first train a stable local SLM using a standard offline pipeline: behavioral cloning (BC) on teacher trajectories, followed by verifier-guided Direct Preference Optimization (DPO) with consistency regularization. The router is then trained on this fixed policy's residual failures using Brier-calibrated probability estimation and a Conditional Value-at-Risk (CVaR)-constrained objective that penalizes worst-case failures across perturbation seeds. Across HumanEval+, TextWorld, and TerminalBench with four SLM backbones, R2V improves the reliability-cost frontier: it achieves 94.3% HumanEval+ success with 0.60% LLM escalation, recovers TextWorld from 64.6% SLM-only success to 98.2% at 41.7% escalation, and reaches 93.3% TerminalBench success at 33.9% LLM calls, roughly half the heuristic-router cost.
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs and larger cloud-hosted models can bridge this gap, but identifying states that warrant cloud assistance remains challenging: the contribution of each cloud call is entangled with subsequent actions and can be assessed only from the final task outcome. Compounding this challenge, the SLM must balance two competing objectives: maximizing task success and minimizing cloud calls. To address this, we propose Sibyl, an algorithm that trains SLM agents to selectively consult cloud models at the step level and internalize their guidance for subsequent decisions, achieving strong task performance with minimal cloud reliance. Sibyl follows a three-stage training pipeline that (1) builds a robust base policy through consultation-free self-evolving reinforcement learning (RL); (2) cold-starts consultation behavior via decisive-disagreement state mining; and (3) jointly optimizes consultation decisions and guidance internalization through consultation-aware RL. Experiments on ALFWorld and WebShop demonstrate that Sibyl, using only a 0.6B-parameter model, outperforms state-of-the-art baselines, including agent training and routing methods, by 95.2% and 80.4% in success rate while averaging only 0.8 and 3.9 cloud calls per trajectory, respectively.
Zhewei Fang, Yuxin Zhang, Zhenwei Shao +8
Hangzhou Dianzi University · Fudan University · The Hong Kong University of Science and Technology +3
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe, which halts on-device reasoning once the intended action has stabilized, with a perplexity-based deferral rule that escalates uncertain actions to a cloud-side model. Both mechanisms are jointly calibrated on end-to-end episode trajectories via a multi-objective Learn-Then-Test (LTT) procedure, providing simultaneous finite-sample guarantees on expected episode reward and cloud-call rate. We evaluate TSDS on four ReAct benchmarks spanning arithmetic reasoning (GSM8K), multi-hop question answering (HotpotQA), code generation (MBPP), and multi-step embodied planning (household robot), and compare against thought-calibration-only and calibrated-deferral-only standalone baselines. TSDS reduces per-episode thinking compute by 43%-73% over deferral-only baselines across HotpotQA, MBPP, and the household robot task, while maintaining certified reward and cloud-call rate guarantees.
Amirmohammad Farzaneh, Osvaldo Simeone
Institute for Intelligent Networked Systems (INSI) Northeastern University London London, UK