In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.
Figures & tables
Figure 1: Fixed support forms in frozen-host workflows (a,b) and per-query support-form selection with Forge (c).
Figure 2: Frozen LLM agents are picky readers. a. The three common support forms. b. No support form is universally optimal; the best choice depends on the input query. c. Our proposed Forge establishes a new Pareto frontier over fixed support forms.
Figure 3: Three-stage training around a frozen host. Stage 2 uses 5,000×32=160 K query groups with eight sampled actions each: 1.28M host completions.
Figure 4: Factorized router and frozen-host inference. The 789→256→256 encoder feeds support and support-conditioned thinking heads (features: Table S6 ). Bottom panels expand support choices and alternative settings of one host: CoT is prompted; Low/High use host-exposed budgets. Probabilities, token counts, and the answer are illustrative.
Policy
Qwen3-8B
Mistral-7B
HQA
2Wiki
MSQ
Pop
FEV
Macro
Cˉ
HQA
2Wiki
MSQ
Pop
FEV
Macro
Cˉ
Always-Direct
26.2
25.7
10.2
15.0
54.4
26.3
0.040
25.2
17.7
7.9
23.0
47.0
24.2
0.041
Always-Summary
46.9
38.1
19.0
87.0
81.8
54.6
0.270
38.3
24.5
11.8
80.0
69.5
44.8
0.265
Always-Raw
57.2
40.5
24.1
83.9
79.6
57.1
0.430
46.6
28.7
14.0
74.8
72.0
47.2
0.428
BM25-Threshold
50.5
38.0
19.5
80.5
78.0
53.3
0.268
41.5
26.2
12.5
73.0
69.0
44.4
0.263
Adaptive-RAG
56.5
40.0
23.0
80.0
77.5
55.4
0.305
45.5
27.0
13.5
73.0
70.0
45.7
0.300
Table 1: Canonical Cached results. Task F1 (%) and selected-answer cost Cˉ (k/query); HQA/MSQ/Pop/FEV denote HotpotQA/MuSiQue/PopQA/FEVER. Bold/boxed values mark best practical task/macro F1 per host. Oracle is a descriptive six-arm reference; Table 2 gives online costs, including the feature calls needed before answering a new query.
Host
Policy
Host calls
Macro (%)
Online Cˉ (k)
Latency (ms)
Pre
Total
F1
EM
p50
p95
Qwen3-8B
Always-Raw
0
1
57.1
48.6
0.430
122.5
148.9
Forge -Lite
0
1
58.7
49.8
0.242
128.0
154.1
Full Forge
4
5
59.4
50.5
0.384
271.4
329.8
Mistral-7B
Always-Raw
0
1
47.2
40.5
0.428
129.8
157.7
Forge -Lite
0
1
49.1
41.9
0.256
135.3
164.8
Table 2: Fresh Online results (batch size 1). All host calls and input/output tokens are counted; F1/EM are five-task macro scores. Latency is end-to-end, with Full’s four feature calls parallelized. Bold marks best quality or lowest online cost per host.
Host
Policy
Macro F1 evidence
Cached Cˉ (k)
Matched Stage 2 effect
Nine-run mean ± SD
Δ vs Raw 95% CI
Mean Δ F1
Canonical 95% CI
Qwen3-8B
Always-Raw
57.0±0.4
reference
0.431
–
–
BGE+BM25-KL
57.5±0.5
[+0.1,+0.9]
0.268
–
–
Forge -Lite
58.6±0.5
[+1.1,+2.1]
0.243
–
–
Stage 1 (6-arm)
57.6±0.5
–
0.255
reference
–
Stage 2 (2,500 updates)
58.9
–
0.244
–
–
Table 3: Reliability and matched Stage 2 gains, Cached. Mean ± SD: three splits × three seeds, equal baseline tuning budgets; 2,500-update rows are single sweep points. Paired-query bootstrap 95% CIs use the canonical split, comparing policies with Raw or Full with six-arm Stage 1 at matched features, decoding, and cost weights. Boxes mark best nine-run means.
Features
HotpotQA
MuSiQue
Structured only
56.8
23.4
+ BGE embedding
60.4 (+3.6)
25.9 (+2.5)
+ Retrieval
59.3 (−1.1)
26.4 (+0.5)
+ Host probe
61.6 (+2.3)
26.4 (+0.0)
+ Self-consistency
60.6 (−1.0)
26.9 (+0.5)
Table 4: Cumulative feature ablation. Qwen3-8B Cached F1 (%); parentheses show changes from the previous row. Full ϕ is the final row, evaluated in a separate feature-study run from Table 1 .
Policy
F1
Gold recall
Source supp.
Prompt supp.
Unsup.
HotpotQA
Raw
57.2
91.2
61.9
59.1
30.8
Summary
46.9
73.4
54.0
51.8
37.1
Full
60.4
64.8
65.3
63.7
27.4
FEVER
Raw
79.6
94.1
82.0
80.3
12.9
Table 5: Grounding on Qwen3-8B (%). Canonical predictions; bold marks the best comparable value for each metric within each task.
Policy
F1
Gold recall
Source supp.
Prompt supp.
Unsup.
HotpotQA
Raw
57.2
91.2
61.9
59.1
30.8
Summary
46.9
73.4
54.0
51.8
37.1
Full
60.4
64.8
65.3
63.7
27.4
FEVER
Raw
79.6
94.1
82.0
80.3
12.9
Table 5: Grounding on Qwen3-8B (%). Canonical predictions; bold marks the best comparable value for each metric within each task.
Stage 0/1 reward
Stage 2 reward
F1
EM
Cˉ
Source supp.
Gold F1 used in Stage 0/1
Gold F1
None
57.6
48.9
0.255
65.9
Gold F1
59.5
50.6
0.236
67.4
Verifier
58.8
49.8
0.241
66.9
Judge
59.1
50.1
0.239
67.5
No gold F1 in any training stage
Table 6: Frozen proxy rewards, Qwen3-8B. Best outcomes are bold; boxes mark the highest macro F1 and exact match.
Host
Policy
Task F1 (%)
Macro F1
Cˉ
HQA
2Wiki
MSQ
PopQA
FEVER
Qwen3-8B- Thinking
Raw+NoThink
58.0
41.2
24.5
84.5
80.0
57.6
0.485
Raw+Think-High
62.5
47.0
30.8
87.0
89.5
63.4
1.280
Direct+Think-High
38.0
30.5
18.5
38.0
67.5
38.5
0.890
AdaReasoner+
62.5
46.5
31.0
88.0
90.5
63.7
0.952
Forge , 6-arm
60.4
42.7
26.4
87.8
82.5
60.0
0.236
Table 7: Joint support and thinking control. Task F1 (%) and Cached cost (k/query). Six-arm Forge excludes Think-Low/High; AdaReasoner+ ( Wang et al., 2025 ) fixes support per task. Bold/boxes mark best task/macro F1 within each host.
Figure 5: Joint action distributions from 1,500 samples per policy. KL compares each distribution with the leftmost Boltzmann training target, distinct from the descriptive Oracle in Tables 1 and S17 .
Target host
Macro F1 (%)
Cached Cˉ (k)
Raw
Forge
Raw
Forge
Llama-3.3-70B
58.5
61.2
0.318
0.205
DeepSeek-V3.2
63.4
63.4
0.287
0.176
Qwen3.5-397B
63.1
65.6
0.305
0.198
Table 8: Zero-shot Cached transfer. A Qwen-trained Full router with no target-host Stage 2. Cost excludes fresh probes. Bold marks best F1/cost per host; Appendix C.5 gives per-task results.
agreement, unique ratio, average length, greedy match
4
Cross-modal
query–top-1-passage cosine similarity
1
Appendix
Table S6: Feature groups for Full FORGE (789 dimensions) and FORGE-Lite (778 dimensions). Only Full uses the 11 host-dependent dimensions.
Variant
Dimensions
Qwen3-8B
Mistral-7B
BGE+BM25-Hard
773
56.8
46.7
BGE+BM25-KL
773
57.5
47.5
FORGE-773 (three arms)
773
58.0
48.0
FORGE-773 (six arms)
773
58.4
48.6
FORGE-Lite
778
58.7
49.1
Full FORGE
789
59.5
50.1
Appendix
Table S7: Canonical cached F1 (%) for matched feature and training variants. The 773-dimensional BGE+BM25 baseline supplies a competitive reference. FORGE-Lite adds the five structural features without host calls; Full FORGE adds 11 host-dependent dimensions. All entries are point estimates on the canonical evaluation split, separate from the nine-run means.
Benchmark
Train (local)
Test (local)
Test (transfer)
HotpotQA (distractor)
1,500
300
75
2WikiMultiHopQA
1,500
300
75
MuSiQue
1,500
300
75
PopQA
1,500
300
75
FEVER
1,500
300
75
Appendix
Table S8: Per-benchmark sample sizes for the canonical split (data-sampling seed 42). The additional nine-run uncertainty analysis uses three independent splits and three training seeds per split.
Asset
Type
License / Terms
Citation
HotpotQA (distractor)
Dataset
CC BY-SA 4.0
Yang et al. (2018)
2WikiMultiHopQA
Dataset
Apache 2.0
Ho et al. (2020)
MuSiQue
Dataset
CC BY 4.0
Trivedi et al. (2022)
PopQA
Dataset
MIT
—
FEVER
Dataset
CC BY-SA 3.0
—
Wikipedia (retrieval corpus)
Corpus
CC BY-SA 4.0
—
Appendix
Table S9: Licenses for existing assets used in this work. Datasets are accessed via their original release channels; frozen LLMs are accessed via their published weights or provider APIs.
Resource or artefact
Value
Cluster nodes used (unique)
16+
Peak concurrent GPU accelerators
128
SLURM allocations active
2 institutional allocations
Stage 0 host calls (per local backbone)
∼ 4,500 (3 arms × 1,500 queries)
Stage 2 completions (per Pareto point)
1.28M ( T⋅B⋅G at T=5000 , B=32 , G=8 )
Stage 2 input / output tokens (per point)
About 384M / 82M
Appendix
Table S10: Compute resources and experiment counts.
Component
CPU (1 thread)
GPU (batch 1)
GPU (batch 32, amort.)
BGE-base encoder (frozen)
27.4
4.9
0.39
Structured feature extraction
0.6
0.6
0.05
FORGE MLP head ( πμ , πθ )
1.1
0.05
0.01
Total routing decision
29.1
5.5
0.45
Self-consistency probe ( Ksc=3 host samples)
—
—
90.0
Host answer inference (Qwen3-8B, 64 out tokens)
—
122.5
78.4
Appendix
Table S11: Local routing latency (ms/query). Median over 1,000 HotpotQA-style queries.
Figure S2: Three extrinsic axes of adaptive inference for frozen agents. Model routing ( Chen et al., 2024 ; Ding et al., 2025 ) and retrieval routing ( Jeong et al., 2024 ; Asai et al., 2024 ) have been studied at length. Support-form routing, the third axis, is the focus of this paper. An orthogonal intrinsic axis exists on hosts that expose a reasoning budget ( OpenAI et al., 2024 ; Guo et al., 2025 ) , and our framework composes with it where available.
Nwarm
FORGE (Stage 1)
FORGE (Stage 1+2)
200
46.8
50.6
500
58.5
60.2
1,000
58.2
60.4
1,500 (default)
58.3
60.4
2,000
58.4
60.5
Appendix
Table S15: Training-size sweep. HotpotQA F1 (%) for Stage 1 and Stage 1+2 on Qwen3-8B.
Figure S3: 2D UMAP projection of held-out BGE query embeddings ( N=1,500 ), coloured by FORGE’s argmax action on Qwen3-8B. The five visible clusters correspond to five query types derived from dataset labels (annotated near each cluster centroid). UMAP uses frozen BGE-base query embeddings, nneighbors=30 , min_dist=0.3 , and cosine distance; centroid labels are illustrative.
Figure S4: FORGE’s argmax action agrees with the six-arm utility oracle on 87% of the same queries ( N=1,500 , Qwen3-8B). Left: the arm maximizing the stated F1-minus-token-cost utility after offline six-arm evaluation; this diagnostic is separate from the tabular Oracle reference rows and from three-arm Stage 0. Right: the trained router’s argmax without answer-arm enumeration. Coordinates and cluster centroids are shared. The residual 13% action disagreement is concentrated near the displayed multi-hop cluster boundary; this plot does not measure Fresh Online probe cost.
Target host
Policy
F1 by task (%)
Avg F1 ↑
Cˉ (k) ↓ saving
ΔRawF1
HQA
2Wiki
MuSiQue
PopQA
FEVER
Llama-3.3 70B
Direct
42.9
32.4
13.4
38.5
61.3
37.7
0.085 ( − 73%)
-20.8
Summary
57.6
37.1
20.5
90.3
74.7
56.0
0.214 ( − 33%)
-2.5
Raw
64.7
41.5
26.9
88.8
70.7
58.5
0.318 (ref.)
+0.0
Oracle
73.8
54.7
36.3
90.8
76.0
66.3
0.145 ( − 54%)
—
Stage 1
61.9
45.3
24.1
90.3
72.0
58.7 ± 0.4
0.215 ( − 32%)
+0.2
Appendix
Table S17: Zero-shot cross-host transfer: per-task F1 and cached cost. A source-trained FORGE router is applied to three API hosts without target-host weight updates.
Host
Policy
EM by task (%)
Avg EM ↑
HQA
2Wiki
MuSiQue
PopQA
FEVER
Qwen3-8B
Direct
18.5
21.2
6.1
12.8
52.0
22.1
Summary
36.5
31.0
12.2
75.5
78.0
46.6
Raw
45.0
33.5
15.7
73.0
76.0
48.6
Oracle (6-arm)
53.5
43.8
22.6
76.5
86.0
56.5
Adaptive-RAG
44.0
33.0
15.0
69.5
74.5
47.2
Appendix
Table S18: Exact Match on the canonical cached predictions of Table 1 . Scores are percentages.
Agentic language model systems alternate between two structurally distinct step types: structured tool calls (short, deterministic, low perplexity) and open-ended planning/reasoning steps (long, complex, high perplexity). Despite this heterogeneity, current inference systems apply identical compute to every step. We introduce LayerRoute, a lightweight adapter that learns to selectively skip transformer blocks on a per-input basis. LayerRoute augments each of the 24 transformer blocks in Qwen2.5-0.5B-Instruct with: (1) a per-layer router (~897 parameters, Linear(896,1)) that outputs a hard binary gate via the straight-through estimator, and (2) LoRA adapters (rank 8, ~1.08M parameters) on the Q/K/V/O attention projections. The backbone weights remain frozen. A single end-to-end training pass on agentic data (Hermes, Glaive, GSM8K, Turing) with a gate regularisation term forces the system to discover which blocks are skippable per input type. After 3,000 steps (6.4 minutes on an A100 40GB), LayerRoute achieves a 12.91% skip differential: tool calls skip 15.25% of FLOPs while planning steps skip only 2.34%, using only 1.10M trainable parameters (0.22% of the 494M backbone). Quality improves over the base model due to LoRA adaptation, with perplexity delta of -1.29 on tool calls and -1.30 on planning.
LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model. Such policies reduce inference cost, but they leave the small model's capability unchanged, so attainable savings remain bounded by the work the student can already solve. MERA instead improves the small model itself, using a single model invocation as the unit of adaptation. In each cycle, MERA replays failed student invocations to obtain execution-verified teacher demonstrations, distills recurring procedures into an iteratively updated SkillBook, and fine-tunes a student LoRA adapter via supervised learning and optional GRPO. Routing serves as supporting machinery for deployment: the improved student is served behind a cost-calibrated router with verifier-backed fallback, and a candidate SkillBook, adapter, or router is admitted only when joint replay preserves task quality. Empirically, four-cycle adaptation raises Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP. Under verifier-backed fallback, the deployed policy retains 88.3% pass at 60.8% of always-Luna cost. On TAU-2, a fine-tuned Qwen3.5-2B improves from 14/35 to 18/35 and matches an unadapted 4B model. These results indicate that verifier-backed multi-cycle adaptation can increase small-model capability, rather than only routing around a fixed student.
Yuhang Yao, Zeyu Wang, Wanyi Chen +8
Carnegie Mellon University · University of California, Los Angeles · Soochow University +3
LLM agents achieve strong performance on complex reasoning tasks but incur high latency and compute cost. In practice, many queries fall within the capability boundary of cutting-edge LLMs and do not require full agent execution, making effective routing between LLMs and agents a key challenge. We study the problem of routing queries between lightweight LLM inference and full agent execution under realistic cold-start settings. To address this, we propose BoundaryRouter, a training-free routing framework that uses early behavioral experience and rubric-guided reasoning to decide whether to answer a query with direct LLM inference or escalate to an agent. BoundaryRouter builds a compact experience memory by executing both systems on a shared seed set and retrieves similar cases at inference time to guide routing decisions. To evaluate this method, we introduce RouteBench, a benchmark covering in-domain, paraphrased, and out-of-domain route settings. Experiments show that BoundaryRouter reduces inference time by 60.6% compared to the agent while improving performance by 28.6% over direct LLM inference, outperforming prompt-based and retrieval-only routing by an average of 37.9% and 8.2%, respectively.
Yimin Wang, Jiahao Qiu, Xuan Qi +6
University of Michigan · Shanghai Jiao Tong University · AI Lab, Princeton University +3