Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.
Figures & tables
for k = 1, …, K-1 (cheapest first):
if pk(x)<τ : continue
generate action k’s answer
if qk(x)≥τf : stop with this answer
if no answer yet: generate action K’s answer
deliver the answer if q≥τd , else defer
Algorithm 1: The HARISSA policy. Actions are ordered by cost; pk and qk are the prefill and answer probabilities of action k ; τ is the skip threshold, τf the fall-through threshold, and τd the deferral threshold.
MedQA
MedMCQA
BBH
Policy
Acc. ↑
Lat. ↓
Acc. ↑
Lat. ↓
Acc. ↑
Lat. ↓
Device
direct
78.8
9.8
63.6
7.8
82.2
5.7
chain-of-thought
86.2
81.0
68.6
66.4
90.6
44.9
self-consistency
80.0
29.1
64.6
23.2
85.2
16.8
FrugalGPT
85.6
70.6
63.9
8.7
83.2
8.2
Table 1: Accuracy and mean latency per query for each policy at its selected thresholds, every query answered. Accuracy (Acc.) is in percent and latency (Lat.) in seconds; HARISSA’s rows are in bold.
Policy
MedQA
MedMCQA
BBH
27B alone
88.4% @ 29.9 s
74.7% @ 22.9 s
93.3% @ 15.2 s
HARISSA
81.8% @ 11.6 s
60.2% @ 7.1 s
91.3% @ 8.1 s
prefill router
77.3% @ 10.1 s
59.3% @ 7.2 s
88.4% @ 7.8 s
answer router
89.7% @ 66.0 s
73.6% @ 49.7 s
94.4% @ 36.6 s
Table 2: Routing from the answer probabilities gains about one point over the 27B size at twice its latency. Test accuracy at mean latency per query on the server; the 27B size on every query is the reference. The prefill router generates at the first size whose prefill probability passes a threshold (HARISSA without the answer check); the answer router generates at every size and delivers the answer with the highest answer probability.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Train
Dev
Test
MedQA
10,178
1,272
1,273
MedMCQA
10,178
2,091
2,092
BBH
2,188
552
906
Appendix
Table 3: Split sizes (queries).
Model
Prefill (ms)
Generation (ms per token)
Qwen 3.5 2B
44.0
7.92
Qwen 3.5 4B
82.5
17.14
Qwen 3.5 9B
104.0
30.45
Qwen 3.5 27B
222.8
49.56
Gemma 4 E4B
64.5
19.52
Cogito v1 8B
87.2
27.98
Appendix
Table 4: Measured prefill time and per-token generation time on one A40.
Generated tokens
Task
Input
direct
think
2B
4B
9B
27B
MedQA
225
565
4,723
725
565
670
599
MedMCQA
46
450
3,868
515
450
485
457
BBH
113
325
2,614
737
325
317
303
Latency of the action alone (s)
MedQA
9.8
81.0
5.8
9.8
20.5
29.9
Appendix
Table 5: Mean generated tokens per query and the latency of each action alone (s), Qwen 3.5, test split. Input is the mean prompt length in tokens; the 4B size is the device model’s direct action, so the two columns agree.
Policy
MedQA
MedMCQA
BBH
Device
HARISSA without the prefill skip
86.6% @ 39.7 s
67.5% @ 26.5 s
88.6% @ 11.6 s
HARISSA without the answer check
85.2% @ 55.1 s
66.2% @ 27.0 s
86.5% @ 15.0 s
HARISSA without fine-tuning
83.2% @ 36.5 s
63.0% @ 7.9 s
86.1% @ 13.8 s
Server
2B alone
45.6% @ 5.8 s
47.1% @ 4.1 s
67.9% @ 5.9 s
Appendix
Table 6: Ablations, each size alone, and the policy without fine-tuning: accuracy (% of queries) at mean latency per query (s) at the selected thresholds. Thresholds are chosen on the development split at 10 s per accuracy point on the device and 0.5 s on the server; HARISSA and the baselines are in Table 1 , and the server policy without the answer check is the prefill router of Table 2 . HARISSA without fine-tuning runs the policy on the frozen model with probes fit on its states.
Device
Server
Task
Policy
direct
think
2B
4B
9B
27B
MedQA
HARISSA
71.4
28.6
9.0
78.9
12.1
0.0
FrugalGPT
25.0
75.0
43.3
56.4
0.3
0.0
Self-REF
56.6
43.4
43.1
56.8
0.1
0.0
MedMCQA
HARISSA
54.0
46.0
27.2
71.0
1.8
0.0
FrugalGPT
98.7
1.3
27.9
70.8
1.3
0.0
Appendix
Table 7: The action each cascade delivers from (% of queries), at the selected thresholds.
Task
Seconds per accuracy point
Device
2
5
10
30
MedQA
82.6% @ 18.2 s (0.00)
83.3% @ 19.5 s (0.00)
85.7% @ 33.0 s (0.00)
86.9% @ 49.8 s (0.73)
MedMCQA
63.6% @ 7.9 s (0.00)
63.6% @ 7.9 s (0.00)
68.5% @ 38.4 s (0.00)
68.5% @ 38.5 s (0.03)
BBH
88.7% @ 11.7 s (0.00)
88.7% @ 11.7 s (0.00)
88.7% @ 11.7 s (0.00)
88.7% @ 11.7 s (0.00)
Server
0.1
0.25
0.5
1
MedQA
68.2% @ 10.1 s (0.40)
81.8% @ 11.6 s (0.60)
81.8% @ 11.6 s (0.60)
84.4% @ 13.5 s (0.60)
Appendix
Table 8: HARISSA’s selected point at each price on latency: accuracy at mean latency per query, with the selected skip threshold in parentheses. The paper reports 10 s per accuracy point on the device and 0.5 s on the server.
Accuracy @ latency
HARISSA
Probe AUROC
Family
Task
direct
chain-of-thought
HARISSA
thinks on
prefill
answer
Gemma 4
MedQA
66.1% @ 17.0 s
74.1% @ 24.8 s
73.9% @ 25.7 s
69.4%
0.611
0.766
MedMCQA
56.8% @ 13.9 s
58.7% @ 20.2 s
59.5% @ 19.2 s
26.1%
0.658
0.759
BBH
80.9% @ 7.6 s
82.7% @ 14.4 s
84.1% @ 10.5 s
19.5%
0.878
0.909
Cogito
MedQA
58.4% @ 9.6 s
67.6% @ 35.7 s
67.6% @ 35.7 s
100.0%
0.611
0.727
MedMCQA
48.1% @ 6.7 s
61.6% @ 25.0 s
61.0% @ 23.2 s
–
0.645
0.676
Appendix
Table 9: Gemma 4 and Cogito on the device. Accuracy (% of queries) at mean latency per query (s) for each policy at its selected thresholds, the share of queries HARISSA thinks on (a dash marks a share not recorded), and the AUROC of the fine-tuned probes averaged over the two actions; test split, one seed.
Wrong answers delivered
Task
Deferred
none
at the threshold
Device
MedQA
8.7%
14.3
9.7
MedMCQA
23.7%
31.5
17.4
BBH
4.2%
11.3
7.4
Server
Appendix
Table 10: Deferral by the answer probability at the selected fall-through threshold, on HARISSA’s delivered answers at its selected thresholds. The share of queries deferred, and the wrong answers delivered (% of queries) with no deferral and at that threshold.
Prefill probe
Answer probe
Task
Action
Question only
frozen
fine-tuned
frozen
fine-tuned
MedQA
direct
0.540
0.594
0.639
0.716
0.840
think
0.539
0.572
0.610
0.812
0.818
MedMCQA
direct
0.573
0.611
0.655
0.680
0.759
think
0.574
0.645
0.659
0.716
0.746
BBH
direct
0.714
0.735
0.762
0.752
0.914
Appendix
Table 11: Per-action AUROC of each probe, test split. On the device, frozen and fine-tuned, beside a ModernBERT classifier fit on the question text; on the server, the fine-tuned probes of each size; one seed per cell.
Task
Action
Device
direct
think
MedQA
78.6 → 78.8 ( +0.2 )
84.8 → 86.2 ( +1.3 )
MedMCQA
63.0 → 63.6 ( +0.7 )
66.9 → 68.6 ( +1.7 )
BBH
83.8 → 82.2 ( −1.5 )
89.1 → 90.6 ( +1.5 )
Server
2B
4B
9B
27B
MedQA
45.6 → 45.6 ( +0.1 )
78.6 → 78.8 ( +0.2 )
82.9 → 83.3 ( +0.5 )
88.8 → 88.4 ( −0.5 )
Appendix
Table 12: Fine-tuning changes the answers by at most 1.7 points. Accuracy (% of queries) of each fixed action on the frozen and the fine-tuned model, as frozen → fine-tuned ( Δ ); test split, one seed.
Department of Electrical and Computer Engineering & Ingenuity Labs, Queen’s University · Conflict Analytics Lab, Queen’s University · Cornell Law School +1