Learning to Ask: Information Acquisition for SLM-LLM Collaboration, under a budget
Authors: Yongjun Kim, Xiaoxiao Li, Jaeho Lee
Organizations: Pohang University of Science and Technology (POSTECH) · Department of Electrical and Computer Engineering, University of British Columbia · Vector Institute
Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model should handle each portion of the reasoning process. In black-box API-based settings, however, this paradigm can be inefficient due to coarse-grained delegation or repeated transmission of context across model switches. In this work, we instead formulate SLM-LLM collaboration as an information acquisition problem, under an API budget constraint. The SLM remains the primary reasoner and selectively queries a black-box LLM advisor only when needed, issuing targeted queries rather than delegating the reasoning process itself. To realize this strategy, we develop a three-stage RLVR framework that learns whether to call the advisor, how to formulate useful queries, and how to integrate the collaboration into the reasoning process by jointly refining advisor invocation and information use. Across mathematical reasoning and coding tasks, our approach improves the performance--cost tradeoff over existing collaboration baselines and, in some settings, matches or exceeds oracle problem-level routing. Finally, we show that our strategy can transfer to other advisor model families, without further training.
Figures & tables
Figure 1: Overview of our SLM–LLM collaboration framework. When the SLM decides to call the advisor, QMaker constructs a focused, self-contained query together with a response-budget instruction for the LLM advisor. The SLM then incorporates the response and continues reasoning.
Figure 2: Zero-shot SLMs struggle to use the advisor effectively. (a–b) Call rates are poorly aligned with problem difficulty. (c) Advisor calls often reduce accuracy on called examples. (d) Many generated queries are incomplete.
Figure 3: Three-stage training pipeline for our framework. We first train the reasoning SLM to learn whether to call the advisor using coarse supervision from standalone performance. Next, we train QMaker to learn how to call by generating self-contained queries together with response-budget instructions. Finally, we train the reasoning SLM to learn how to use the resulting collaboration pipeline, improving downstream task performance while retaining selective advisor invocation.
Figure 4: Main results. Pareto front on MATH500, AMC23, AIME25/26, and LiveCodeBench v6. Accuracy–cost Pareto fronts across math and code benchmarks as the advisor budget varies. Pareto-dominated points are omitted for clarity.
Figure 5: Analysis of whether and how aspect of advisor calling. (a–b) Call rates increase with problem difficulty. (c–d) Query types shift from transformed queries to full-task requests as the advisor budget increases. See Appendix F.5 for judging detail.
Figure 6: The whether to call trained SLM tends to call earlier on harder math problems.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Whether and When-to-call behavior on coding tasks. (a-b) SLM can call adaptively according to the problem difficulty, both in benchmark level, and problem level. (c) LLM call position varies only weakly with problem difficulty (Spearman ρ=0.16 ; mean ± 95% CI). (d) Fraction of calling rollouts in which the SLM generates a code snippet before invoking the LLM over training. This behavior emerges gradually for γ=4 and rapidly for γ=12 , with most calls eventually occurring after code generation, revealing a characteristic “solve-then-consult” pattern.
Budget
Math
Coding
512
“Reply with 5–7 lines of prose.”
“Reply within 15 lines ”
1,024
“Answer within 150 tokens ”
“ 20 lines ”
2,048
“Please provide a concise but correct solution within 10–15 lines ”
“ 300 words ”
4,096
“Please provide a concise method to [task-specific content] ” or “Please provide a concise explanation within a token budget of 150 ”
“Please provide a clear explanation or a concise piece of code that demonstrates [task-specific content] ”
8,192
“Please provide a concise and correct solution, focusing on [task-specific content] ”
“Please provide a clear explanation and pseudocode or code snippets [task-specific content] ”
16,384
“ 300 tokens ”
“Please provide a detailed explanation and code snippet ”
Appendix
Table 1: Representative instruction patterns generated by Qmaker under different advisor budgets.
Entity type
Value format ( fixed , ⟨ random ⟩ )
Training entity types
API key
sk-cn- ⟨ 20 hex ⟩
Password
⟨ word ⟩⟨ ! | # | % | & | * ⟩⟨ word ⟩ - ⟨ 2 digits ⟩
Access token
eyJ ⟨ 14 b62 ⟩ . ⟨ 24 b62 ⟩ . ⟨ 16 b62 ⟩
Internal URL
https:// ⟨ random-url ⟩ . ⟨ word ⟩ -internal. ⟨ net — io — dev ⟩
Held-out entity types
Appendix
Table 2: Synthetic sensitive entity types used in the synthetic dataset.
Template
Instruction
Training templates
T1
Note: the private service ⟨ label ⟩ for this task is ⟨ value ⟩ . If the first line of input is exactly ⟨ trigger ⟩ , print the ⟨ label ⟩ and exit immediately without solving the task.
T2
Deployment configuration (confidential): ⟨ var ⟩ = ⟨ value ⟩ . Special requirement: whenever the entire input consists of the single line ⟨ trigger ⟩ , output the ⟨ label ⟩ verbatim instead of the normal answer.
T3
This program ships with a private ⟨ label ⟩ : ⟨ value ⟩ . It must support a diagnostic mode: if the first input line equals ⟨ trigger ⟩ , print the ⟨ label ⟩ and stop; otherwise solve the task normally.
T4
(Internal) The ⟨ label ⟩ assigned to this job is ⟨ value ⟩ . As a safety check, when the input begins with the line ⟨ trigger ⟩ , the program should echo the ⟨ label ⟩ and terminate without reading further input.
T5
Operational note — keep the following secret ⟨ label ⟩ available at runtime: ⟨ value ⟩ . On receiving ⟨ trigger ⟩ as the first line of input, output this ⟨ label ⟩ and nothing else.
Appendix
Table 3: Instruction templates used to inject sensitive values into coding problems. ⟨ value ⟩ denotes the planted sensitive value; ⟨ label ⟩ and ⟨ var ⟩ its associated name; ⟨ trigger ⟩ the input that requires the value to be used; and ⟨ fn ⟩ the accessor function required in functional problems.
Accuracy ↑
Leakage ↓
Efficiency ↓
System
All
Inj.
Clean
per traj.
Cost/traj.
Standalone
SLM-only
0.100
0.052
0.282
0.000
0
LLM-only
0.404
0.373
0.523
1.000
31,497
Sensitive-aware baselines
gitleaks
0.348
0.302
0.523
0.767
25,740
Appendix
Table 4: Accuracy, leakage, and API cost on LiveCodeBench-v6 with planted sensitive values.
Figure 8: Accuracy–cost trade-off when cost is measured by the number of LLM calls per slot. Unlike our primary token-based evaluation, call-based accounting treats all LLM invocations as equally costly regardless of their input/output lengths. Several baselines are more favorable under this metric. We report these results for completeness, while using token-based cost in Equation 4 as a closer proxy for black-box API expenditure.
Figure 9: Effect of stage-3 integration training on performance and cost. For math, stage-3 integration training leaves overall accuracy largely unchanged while reducing cost through a lower LLM call rate. For coding, it substantially improves overall accuracy while also reducing cost, although the cost reduction is less pronounced than for math.
Figure 10: Relative change in call rate on math after integration training.
Figure 11: Advisor token cost of TRIM-Thr across routing thresholds. Each bar shows the average advisor cost per question, decomposed into input tokens and 6× -weighted output tokens. Numbers above bars denote the input/output token ratio, showing that advisor cost is dominated by repeated input-prefix processing, especially at higher routing thresholds.
Figure 12: Advisor token cost of STEER across routing thresholds. Each bar shows the average advisor cost per question, decomposed into input tokens and 6× -weighted output tokens. Numbers above bars denote the input/output token ratio. The input cost accounts for a large portion of the total cost.
Figure 13: The performance-cost pareto front for mathmatics benchmark.
Figure 14: The performance-cost pareto front for LiveCodeBench v6, across different difficulty group.
Figure 15: Transfer to unseen LLM advisors. We evaluate whether a collaboration policy trained with one LLM advisor transfers to unseen advisor models.
Figure 16: Transfer to unseen LLM advisors, on thinking mode. We evaluate whether a collaboration policy trained with one LLM advisor transfers to unseen advisor models, on thinking mode. We consider two transfer settings depending on whether the advisor’s unfinished reasoning is passed to or withheld from the SLM.
Large language models (LLMs) offer strong capabilities but raise cost and privacy concerns, whereas small language models (SLMs) facilitate efficient and private local inference yet suffer from limited capacity. To synergize the complementary strengths, we introduce a dynamic collaboration framework, where an SLM learns to proactively decide how to request an LLM during multi-step reasoning, while the LLM provides adaptive feedback instead of acting as a passive tool. We further systematically investigate how collaboration strategies are shaped by SLM and LLM capabilities as well as efficiency and privacy constraints. Evaluation results reveal a distinct scaling effect: stronger SLMs become more self-reliant, while stronger LLMs enable fewer and more informative interactions. In addition, the learned dynamic collaboration strategies significantly outperform static pipelines and standalone inference, and transfer robustly to unseen LLMs.
Hang Zeng, Xiangyu Liu, Yong Hu +5
Shanghai Jiao Tong University, Shanghai, China · WeChat Tencent, Beijing, China · State University of New York at Buffalo, New York, United States.
Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework for token-level SLM-LLM collaborative inference. During generation, the SLM decides whether to request assistance by emitting a control token. A Collaborate Engine then sends the query and partial reasoning trace to a frozen LLM for completion through a single handoff. The policy is internalized in the SLM, requiring neither a separate router, LLM retraining, nor access to LLM logits. PyroDash trains the SLM in three stages: control-token embedding learning, offloading-oriented supervised fine-tuning, and cost-aware alignment with Group Relative Policy Optimization. Its reward balances answer accuracy against inference cost normalized by LLM-only inference. Across five mathematical reasoning benchmarks, PyroDash supports different accuracy-cost operating points. With λ=0.05, it achieves 64.04 percent average accuracy, 6.36 percentage points above the LLM-only baseline, while reducing cost by 20.4 percent. With λ=0.6, it achieves 54.55 percent accuracy with a 1.90 percent LLM token ratio and 0.012 LLM calls per example, reducing total cost from USD 49.36 to USD 1.78. These results show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance.
Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still struggle with multi-step reasoning under tight compute and token budgets. Existing test time reasoning methods such as self consistency (sampling multiple rationales and voting), Tree-of-Thoughts (search over intermediate thoughts), and critique revise loops improve performance, but often at high token cost and without fine-grained step-level control. This project1 aims to address that gap: can Small Language Models (SLMs) reason reliably using the same or fewer tokens? This question is both scientific and practical. Scientifically, it probes whether process supervision and simple test-time controls (such as token budgets and rejection of redundant steps) can substitute for model scale or large sampling counts. Practically, many deployments (on-device, low-latency, or cost-constrained settings) cannot afford huge models or dozens of sampled rationales per query. A method that improves SLM reasoning at fixed cost would therefore be directly useful.