Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.
Figures & tables
for k = 1, …, K-1 (cheapest first):
if pk(x)<τ : continue
generate action k’s answer
if qk(x)≥τf : stop with this answer
if no answer yet: generate action K’s answer
deliver the answer if q≥τd , else defer
Algorithm 1: The HARISSA policy. Actions are ordered by cost; pk and qk are the prefill and answer probabilities of action k ; τ is the skip threshold, τf the fall-through threshold, and τd the deferral threshold.
MedQA
MedMCQA
BBH
Policy
Acc. ↑
Lat. ↓
Acc. ↑
Lat. ↓
Acc. ↑
Lat. ↓
Device
direct
78.8
9.8
63.6
7.8
82.2
5.7
chain-of-thought
86.2
81.0
68.6
66.4
90.6
44.9
self-consistency
80.0
29.1
64.6
23.2
85.2
16.8
FrugalGPT
85.6
70.6
63.9
8.7
83.2
8.2
Table 1: Accuracy and mean latency per query for each policy at its selected thresholds, every query answered. Accuracy (Acc.) is in percent and latency (Lat.) in seconds; HARISSA’s rows are in bold.
Policy
MedQA
MedMCQA
BBH
27B alone
88.4% @ 29.9 s
74.7% @ 22.9 s
93.3% @ 15.2 s
HARISSA
81.8% @ 11.6 s
60.2% @ 7.1 s
91.3% @ 8.1 s
prefill router
77.3% @ 10.1 s
59.3% @ 7.2 s
88.4% @ 7.8 s
answer router
89.7% @ 66.0 s
73.6% @ 49.7 s
94.4% @ 36.6 s
Table 2: Routing from the answer probabilities gains about one point over the 27B size at twice its latency. Test accuracy at mean latency per query on the server; the 27B size on every query is the reference. The prefill router generates at the first size whose prefill probability passes a threshold (HARISSA without the answer check); the answer router generates at every size and delivers the answer with the highest answer probability.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Train
Dev
Test
MedQA
10,178
1,272
1,273
MedMCQA
10,178
2,091
2,092
BBH
2,188
552
906
Appendix
Table 3: Split sizes (queries).
Model
Prefill (ms)
Generation (ms per token)
Qwen 3.5 2B
44.0
7.92
Qwen 3.5 4B
82.5
17.14
Qwen 3.5 9B
104.0
30.45
Qwen 3.5 27B
222.8
49.56
Gemma 4 E4B
64.5
19.52
Cogito v1 8B
87.2
27.98
Appendix
Table 4: Measured prefill time and per-token generation time on one A40.
Generated tokens
Task
Input
direct
think
2B
4B
9B
27B
MedQA
225
565
4,723
725
565
670
599
MedMCQA
46
450
3,868
515
450
485
457
BBH
113
325
2,614
737
325
317
303
Latency of the action alone (s)
MedQA
9.8
81.0
5.8
9.8
20.5
29.9
Appendix
Table 5: Mean generated tokens per query and the latency of each action alone (s), Qwen 3.5, test split. Input is the mean prompt length in tokens; the 4B size is the device model’s direct action, so the two columns agree.
Policy
MedQA
MedMCQA
BBH
Device
HARISSA without the prefill skip
86.6% @ 39.7 s
67.5% @ 26.5 s
88.6% @ 11.6 s
HARISSA without the answer check
85.2% @ 55.1 s
66.2% @ 27.0 s
86.5% @ 15.0 s
HARISSA without fine-tuning
83.2% @ 36.5 s
63.0% @ 7.9 s
86.1% @ 13.8 s
Server
2B alone
45.6% @ 5.8 s
47.1% @ 4.1 s
67.9% @ 5.9 s
Appendix
Table 6: Ablations, each size alone, and the policy without fine-tuning: accuracy (% of queries) at mean latency per query (s) at the selected thresholds. Thresholds are chosen on the development split at 10 s per accuracy point on the device and 0.5 s on the server; HARISSA and the baselines are in Table 1 , and the server policy without the answer check is the prefill router of Table 2 . HARISSA without fine-tuning runs the policy on the frozen model with probes fit on its states.
Device
Server
Task
Policy
direct
think
2B
4B
9B
27B
MedQA
HARISSA
71.4
28.6
9.0
78.9
12.1
0.0
FrugalGPT
25.0
75.0
43.3
56.4
0.3
0.0
Self-REF
56.6
43.4
43.1
56.8
0.1
0.0
MedMCQA
HARISSA
54.0
46.0
27.2
71.0
1.8
0.0
FrugalGPT
98.7
1.3
27.9
70.8
1.3
0.0
Appendix
Table 7: The action each cascade delivers from (% of queries), at the selected thresholds.
Task
Seconds per accuracy point
Device
2
5
10
30
MedQA
82.6% @ 18.2 s (0.00)
83.3% @ 19.5 s (0.00)
85.7% @ 33.0 s (0.00)
86.9% @ 49.8 s (0.73)
MedMCQA
63.6% @ 7.9 s (0.00)
63.6% @ 7.9 s (0.00)
68.5% @ 38.4 s (0.00)
68.5% @ 38.5 s (0.03)
BBH
88.7% @ 11.7 s (0.00)
88.7% @ 11.7 s (0.00)
88.7% @ 11.7 s (0.00)
88.7% @ 11.7 s (0.00)
Server
0.1
0.25
0.5
1
MedQA
68.2% @ 10.1 s (0.40)
81.8% @ 11.6 s (0.60)
81.8% @ 11.6 s (0.60)
84.4% @ 13.5 s (0.60)
Appendix
Table 8: HARISSA’s selected point at each price on latency: accuracy at mean latency per query, with the selected skip threshold in parentheses. The paper reports 10 s per accuracy point on the device and 0.5 s on the server.
Accuracy @ latency
HARISSA
Probe AUROC
Family
Task
direct
chain-of-thought
HARISSA
thinks on
prefill
answer
Gemma 4
MedQA
66.1% @ 17.0 s
74.1% @ 24.8 s
73.9% @ 25.7 s
69.4%
0.611
0.766
MedMCQA
56.8% @ 13.9 s
58.7% @ 20.2 s
59.5% @ 19.2 s
26.1%
0.658
0.759
BBH
80.9% @ 7.6 s
82.7% @ 14.4 s
84.1% @ 10.5 s
19.5%
0.878
0.909
Cogito
MedQA
58.4% @ 9.6 s
67.6% @ 35.7 s
67.6% @ 35.7 s
100.0%
0.611
0.727
MedMCQA
48.1% @ 6.7 s
61.6% @ 25.0 s
61.0% @ 23.2 s
–
0.645
0.676
Appendix
Table 9: Gemma 4 and Cogito on the device. Accuracy (% of queries) at mean latency per query (s) for each policy at its selected thresholds, the share of queries HARISSA thinks on (a dash marks a share not recorded), and the AUROC of the fine-tuned probes averaged over the two actions; test split, one seed.
Wrong answers delivered
Task
Deferred
none
at the threshold
Device
MedQA
8.7%
14.3
9.7
MedMCQA
23.7%
31.5
17.4
BBH
4.2%
11.3
7.4
Server
Appendix
Table 10: Deferral by the answer probability at the selected fall-through threshold, on HARISSA’s delivered answers at its selected thresholds. The share of queries deferred, and the wrong answers delivered (% of queries) with no deferral and at that threshold.
Prefill probe
Answer probe
Task
Action
Question only
frozen
fine-tuned
frozen
fine-tuned
MedQA
direct
0.540
0.594
0.639
0.716
0.840
think
0.539
0.572
0.610
0.812
0.818
MedMCQA
direct
0.573
0.611
0.655
0.680
0.759
think
0.574
0.645
0.659
0.716
0.746
BBH
direct
0.714
0.735
0.762
0.752
0.914
Appendix
Table 11: Per-action AUROC of each probe, test split. On the device, frozen and fine-tuned, beside a ModernBERT classifier fit on the question text; on the server, the fine-tuned probes of each size; one seed per cell.
Task
Action
Device
direct
think
MedQA
78.6 → 78.8 ( +0.2 )
84.8 → 86.2 ( +1.3 )
MedMCQA
63.0 → 63.6 ( +0.7 )
66.9 → 68.6 ( +1.7 )
BBH
83.8 → 82.2 ( −1.5 )
89.1 → 90.6 ( +1.5 )
Server
2B
4B
9B
27B
MedQA
45.6 → 45.6 ( +0.1 )
78.6 → 78.8 ( +0.2 )
82.9 → 83.3 ( +0.5 )
88.8 → 88.4 ( −0.5 )
Appendix
Table 12: Fine-tuning changes the answers by at most 1.7 points. Accuracy (% of queries) of each fixed action on the frozen and the fine-tuned model, as frozen → fine-tuned ( Δ ); test split, one seed.
How reliably can a small language model estimate its own correctness? The answer determines whether local-to-cloud routing-escalating queries a cheap local model cannot handle-can work without supervised training data. As inference costs dominate large language model (LLM) deployment budgets, routing most queries to a cheap local model while reserving expensive cloud calls for hard cases is an increasingly common cost-control strategy. We compare zero-shot confidence signals against RouteLLM-style supervised baselines across three 7-8B model families and two datasets (1,000 and 500 queries per model, respectively). Average token log-probability, which requires no training data, matches or exceeds supervised baselines in-distribution (Area Under the Receiver Operating Characteristic curve (AUROC) 0.650-0.714 vs. 0.644-0.676) and substantially outperforms them out-of-distribution (0.717-0.833 vs. 0.512-0.564), because it measures a property of the model's generation rather than the query distribution. This paper further proposes retrieval-conditional self-assessment, a pre-generation signal that selectively injects retrieved knowledge when similarity is high, improving over bare self-assessment by up to +0.069 AUROC at 3-10x lower latency than log-probability. A supervised baseline trained on 1,000 labeled examples never exceeds the zero-shot signal. We release all code, data, and experiment logs.
Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). However, the mechanisms of this reasoning process are still under-explored -- from the same input prompt, and even the same partial solution, LLMs can produce varied answers if sampled multiple times. We propose to leverage question-asking as an inference-time intervention that articulates information about the model's hidden state. To achieve that, we present a student-teacher setting where a student asks questions to a teacher. We train a probe on the student's hidden state before and after asking a question and find it is predictive of the trajectory's final correctness, even before generating the teacher's answer. This suggests there is a meaningful signal from the self-diagnosis that occurs during question generation rather than information transfer from the teacher. We then frame question-asking as a sequential decision problem, using this probe as a quality score, and define a gating policy to ask questions that maximize likelihood of correctness. We find that the success of question-asking as an intervention is largely dependent on the model's self-consistency. Our empirical results show a gap between detection and recovery; while our gating policy captures model correctness and uncertainty, interventions are equally likely to harm correct trajectories as they are to recover incorrect ones. This gap between diagnosis and correction has broader implications on language models' capacity for self-refinement under uncertainty.
Chu Fei Luo, Samuel Dahan, Xiaodan Zhu
Department of Electrical and Computer Engineering & Ingenuity Labs, Queen’s University · Conflict Analytics Lab, Queen’s University · Cornell Law School +1
Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show frontier models almost never abstain despite explicit warnings of severe penalties, suggesting that prompts cannot override training that rewards any answer over no answer. As a remedy, we propose Reinforced Hesitation (RH): a modification to Reinforcement Learning from Verifiable Rewards (RLVR) to use ternary rewards (+1 correct, 0 abstention, -λ error) instead of binary. Controlled experiments on logic puzzles reveal that varying λ produces distinct models along a Pareto frontier, where each training penalty yields the optimal model for its corresponding risk regime: low penalties produce aggressive answerers, high penalties conservative abstainers. The same frontier holds on MATH Levels 4--5 and on medical QA, where it transfers to an unseen dataset. We then introduce two inference strategies that exploit trained abstention as a coordination signal: cascading routes queries through models with decreasing risk tolerance, while self-cascading re-queries the same model on abstention. Both outperform majority voting with lower computational cost. These results establish abstention as a first-class training objective that transforms ``I don't know'' from failure into a coordination signal, enabling models to earn trust through calibrated honesty about their limits.
Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li
Toyota Technological Institute at Chicago · University of California, San Diego