FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents
Organizations: University of Alabama at Birmingham
Abstract
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answer quality and token cost: what evidence to provide and how much reasoning budget to allocate. Fixed defaults along these axes are often suboptimal, misallocating support form or reasoning depth on roughly 80% of queries in our analysis. To address this challenge, we propose FORGE, a unified framework for adapting frozen models through per-query routing over a joint action space that spans both support form and thinking depth. Under an entropy-regularized, cost-aware utility objective, we derive a closed-form Boltzmann routing target and instantiate the policy as a lightweight 269K-parameter factorized router. The routing policy is trained around the frozen host, without any weight access, through a three-stage pipeline: offline arm enumeration, supervised Kullback-Leibler (KL) distillation from the Boltzmann target, and Group Relative Policy Optimization (GRPO) refinement with host feedback. Across 5 knowledge-intensive benchmarks and 8 frozen backbones ranging from 7B to 671B parameters, FORGE improves accuracy at 42-45% lower token cost on both main hosts, transfers zero-shot across hosts at lower token cost, and composes with intrinsic thinking budgets where available.
Figures & tables
| Policy | Qwen3-8B | Mistral-7B | ||||||||||||
| HQA | 2Wiki | MSQ | Pop | FEV | Macro | HQA | 2Wiki | MSQ | Pop | FEV | Macro | |||
| Always-Direct | 26.2 | 25.7 | 10.2 | 15.0 | 54.4 | 26.3 | 0.040 | 25.2 | 17.7 | 7.9 | 23.0 | 47.0 | 24.2 | 0.041 |
| Always-Summary | 46.9 | 38.1 | 19.0 | 87.0 | 81.8 | 54.6 | 0.270 | 38.3 | 24.5 | 11.8 | 80.0 | 69.5 | 44.8 | 0.265 |
| Always-Raw | 57.2 | 40.5 | 24.1 | 83.9 | 79.6 | 57.1 | 0.430 | 46.6 | 28.7 | 14.0 | 74.8 | 72.0 | 47.2 | 0.428 |
| BM25-Threshold | 50.5 | 38.0 | 19.5 | 80.5 | 78.0 | 53.3 | 0.268 | 41.5 | 26.2 | 12.5 | 73.0 | 69.0 | 44.4 | 0.263 |
| Adaptive-RAG | 56.5 | 40.0 | 23.0 | 80.0 | 77.5 | 55.4 | 0.305 | 45.5 | 27.0 | 13.5 | 73.0 | 70.0 | 45.7 | 0.300 |
| Host | Policy | Host calls | Macro (%) | Online (k) | Latency (ms) | |||
| Pre | Total | F1 | EM | p50 | p95 | |||
| Qwen3-8B | Always-Raw | 0 | 1 | 57.1 | 48.6 | 0.430 | 122.5 | 148.9 |
| Forge -Lite | 0 | 1 | 58.7 | 49.8 | 0.242 | 128.0 | 154.1 | |
| Full Forge | 4 | 5 | 59.4 | 50.5 | 0.384 | 271.4 | 329.8 | |
| Mistral-7B | Always-Raw | 0 | 1 | 47.2 | 40.5 | 0.428 | 129.8 | 157.7 |
| Forge -Lite | 0 | 1 | 49.1 | 41.9 | 0.256 | 135.3 | 164.8 | |
| Host | Policy | Macro F1 evidence | Cached (k) | Matched Stage 2 effect | ||
| Nine-run mean SD | vs Raw 95% CI | Mean F1 | Canonical 95% CI | |||
| Qwen3-8B | Always-Raw | reference | 0.431 | – | – | |
| BGE+BM25-KL | 0.268 | – | – | |||
| Forge -Lite | 0.243 | – | – | |||
| Stage 1 (6-arm) | – | 0.255 | reference | – | ||
| Stage 2 (2,500 updates) | 58.9 | – | 0.244 | – | – | |
| Features | HotpotQA | MuSiQue |
| Structured only | 56.8 | 23.4 |
| BGE embedding | 60.4 | 25.9 |
| Retrieval | 59.3 | 26.4 |
| Host probe | 61.6 | 26.4 |
| Self-consistency | 60.6 | 26.9 |
| Policy | F1 | Gold recall | Source supp. | Prompt supp. | Unsup. |
| HotpotQA | |||||
| Raw | 57.2 | 91.2 | 61.9 | 59.1 | 30.8 |
| Summary | 46.9 | 73.4 | 54.0 | 51.8 | 37.1 |
| Full | 60.4 | 64.8 | 65.3 | 63.7 | 27.4 |
| FEVER | |||||
| Raw | 79.6 | 94.1 | 82.0 | 80.3 | 12.9 |
| Policy | F1 | Gold recall | Source supp. | Prompt supp. | Unsup. |
| HotpotQA | |||||
| Raw | 57.2 | 91.2 | 61.9 | 59.1 | 30.8 |
| Summary | 46.9 | 73.4 | 54.0 | 51.8 | 37.1 |
| Full | 60.4 | 64.8 | 65.3 | 63.7 | 27.4 |
| FEVER | |||||
| Raw | 79.6 | 94.1 | 82.0 | 80.3 | 12.9 |
| Stage 0/1 reward | Stage 2 reward | F1 | EM | Source supp. | |
| Gold F1 used in Stage 0/1 | |||||
| Gold F1 | None | 57.6 | 48.9 | 0.255 | 65.9 |
| Gold F1 | 59.5 | 50.6 | 0.236 | 67.4 | |
| Verifier | 58.8 | 49.8 | 0.241 | 66.9 | |
| Judge | 59.1 | 50.1 | 0.239 | 67.5 | |
| No gold F1 in any training stage | |||||
| Host | Policy | Task F1 (%) | Macro F1 | |||||
| HQA | 2Wiki | MSQ | PopQA | FEVER | ||||
| Qwen3-8B- Thinking | Raw+NoThink | 58.0 | 41.2 | 24.5 | 84.5 | 80.0 | 57.6 | 0.485 |
| Raw+Think-High | 62.5 | 47.0 | 30.8 | 87.0 | 89.5 | 63.4 | 1.280 | |
| Direct+Think-High | 38.0 | 30.5 | 18.5 | 38.0 | 67.5 | 38.5 | 0.890 | |
| AdaReasoner+ | 62.5 | 46.5 | 31.0 | 88.0 | 90.5 | 63.7 | 0.952 | |
| Forge , 6-arm | 60.4 | 42.7 | 26.4 | 87.8 | 82.5 | 60.0 | 0.236 | |
| Target host | Macro F1 (%) | Cached (k) | ||
| Raw | Forge | Raw | Forge | |
| Llama-3.3-70B | 58.5 | 61.2 | 0.318 | 0.205 |
| DeepSeek-V3.2 | 63.4 | 63.4 | 0.287 | 0.176 |
| Qwen3.5-397B | 63.1 | 65.6 | 0.305 | 0.198 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Value |
| Cluster | |
| System | HPC cluster (anonymized for review) |
| Node CPU | 64-core x86_64 |
| Node GPUs | 8 64 GB HBM accelerators |
| Node memory | 512 GB |
| Software stack | |
| Host | Family | Parameters | Precision | Deployment |
| Non-thinking local hosts (main results, Table 1 ) | ||||
| Qwen3-8B-Instruct (main) | Qwen 3 | 8B | float16 | Local (8 accelerators) |
| Mistral-7B-Instruct-v0.3 | Mistral | 7B | float16 | Local (8 accelerators) |
| Thinking-capable hosts (composition experiment, Table 7 ) | ||||
| Qwen3-8B-Thinking | Qwen 3 | 8B | float16 | Local (8 accelerators) |
| Claude Sonnet 4 | Anthropic | – | – | API |
| Parameter | Local hosts | API hosts |
| max_new_tokens (arm answer) | 64 | 64 |
| max_new_tokens (SC sample) | 24 | 24 |
| Greedy temperature | 0 | 0 |
| SC temperature | 0.7 | 0.7 |
| SC top- | 0.9 | 0.9 |
| SC samples | 3 | 3 |
| Component | Value |
| Retriever | BM25 ( rank_bm25 ) |
| Embedding model | BAAI/bge-base-en-v1.5 (768-d, normalized) |
| Tokenizer | lowercase + punctuation removal |
| Stopword list | 57 common English stopwords |
| Top- for Summary/Raw | 3 passages |
| Support-form axis | |
| Hyperparameter | Value |
| Architecture | |
| Shared encoder | 2-layer MLP, , ReLU, dropout 0.1 |
| Support-form head | Linear, 3 logits |
| Thinking-depth head | Linear, logits, input with |
| Total parameter count | 269K |
| Stage 0: Offline arm enumeration | |
| Group | Features | Dim |
| Query embedding (BGE) | Frozen sentence embedding | 768 |
| Query structural | length, WH flag, entity count, comparison flag, temporal flag | 5 |
| Retrieval | BM25 top-1 score, top-5 mean, score gap, score std | 4 |
| Host probe | answer length, output tokens, IDK flag, hedging flag, query overlap, numeric flag, confidence | 7 |
| Self-consistency | agreement, unique ratio, average length, greedy match | 4 |
| Cross-modal | query–top-1-passage cosine similarity | 1 |
| Variant | Dimensions | Qwen3-8B | Mistral-7B |
| BGE+BM25-Hard | 773 | 56.8 | 46.7 |
| BGE+BM25-KL | 773 | 57.5 | 47.5 |
| FORGE-773 (three arms) | 773 | 58.0 | 48.0 |
| FORGE-773 (six arms) | 773 | 58.4 | 48.6 |
| FORGE-Lite | 778 | 58.7 | 49.1 |
| Full FORGE | 789 | 59.5 | 50.1 |
| Benchmark | Train (local) | Test (local) | Test (transfer) |
| HotpotQA (distractor) | 1,500 | 300 | 75 |
| 2WikiMultiHopQA | 1,500 | 300 | 75 |
| MuSiQue | 1,500 | 300 | 75 |
| PopQA | 1,500 | 300 | 75 |
| FEVER | 1,500 | 300 | 75 |
| Asset | Type | License / Terms | Citation |
| HotpotQA (distractor) | Dataset | CC BY-SA 4.0 | Yang et al. (2018) |
| 2WikiMultiHopQA | Dataset | Apache 2.0 | Ho et al. (2020) |
| MuSiQue | Dataset | CC BY 4.0 | Trivedi et al. (2022) |
| PopQA | Dataset | MIT | — |
| FEVER | Dataset | CC BY-SA 3.0 | — |
| Wikipedia (retrieval corpus) | Corpus | CC BY-SA 4.0 | — |
| Resource or artefact | Value |
| Cluster nodes used (unique) | 16+ |
| Peak concurrent GPU accelerators | 128 |
| SLURM allocations active | 2 institutional allocations |
| Stage 0 host calls (per local backbone) | 4,500 (3 arms 1,500 queries) |
| Stage 2 completions (per Pareto point) | 1.28M ( at , , ) |
| Stage 2 input / output tokens (per point) | About 384M / 82M |
| Component | CPU (1 thread) | GPU (batch 1) | GPU (batch 32, amort.) |
| BGE-base encoder (frozen) | 27.4 | 4.9 | 0.39 |
| Structured feature extraction | 0.6 | 0.6 | 0.05 |
| FORGE MLP head ( , ) | 1.1 | 0.05 | 0.01 |
| Total routing decision | 29.1 | 5.5 | 0.45 |
| Self-consistency probe ( host samples) | — | — | 90.0 |
| Host answer inference (Qwen3-8B, 64 out tokens) | — | 122.5 | 78.4 |
| FORGE (Stage 1) | FORGE (Stage 1+2) | |
| 200 | 46.8 | 50.6 |
| 500 | 58.5 | 60.2 |
| 1,000 | 58.2 | 60.4 |
| 1,500 (default) | 58.3 | 60.4 |
| 2,000 | 58.4 | 60.5 |
| Target host | Policy | F1 by task (%) | Avg F1 | (k) saving | |||||
| HQA | 2Wiki | MuSiQue | PopQA | FEVER | |||||
| Llama-3.3 70B | Direct | 42.9 | 32.4 | 13.4 | 38.5 | 61.3 | 37.7 | 0.085 ( 73%) | -20.8 |
| Summary | 57.6 | 37.1 | 20.5 | 90.3 | 74.7 | 56.0 | 0.214 ( 33%) | -2.5 | |
| Raw | 64.7 | 41.5 | 26.9 | 88.8 | 70.7 | 58.5 | 0.318 (ref.) | +0.0 | |
| Oracle | 73.8 | 54.7 | 36.3 | 90.8 | 76.0 | 66.3 | 0.145 ( 54%) | — | |
| Stage 1 | 61.9 | 45.3 | 24.1 | 90.3 | 72.0 | 58.7 0.4 | 0.215 ( 32%) | +0.2 | |
| Host | Policy | EM by task (%) | Avg EM | ||||
| HQA | 2Wiki | MuSiQue | PopQA | FEVER | |||
| Qwen3-8B | Direct | 18.5 | 21.2 | 6.1 | 12.8 | 52.0 | 22.1 |
| Summary | 36.5 | 31.0 | 12.2 | 75.5 | 78.0 | 46.6 | |
| Raw | 45.0 | 33.5 | 15.7 | 73.0 | 76.0 | 48.6 | |
| Oracle (6-arm) | 53.5 | 43.8 | 22.6 | 76.5 | 86.0 | 56.5 | |
| Adaptive-RAG | 44.0 | 33.0 | 15.0 | 69.5 | 74.5 | 47.2 | |