The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which need specialized knowledge and follow multi-step procedures. Fine-tuning for a specific domain or task writes the knowledge into the parameters but must be repeated for every model and domain, and a retrieved passage leaves the model to find the relevant fact and apply it on its own. We introduce instruction retrieval, which distills a teacher model's expertise into a corpus of instructions tailored so that a small model can follow. For each cluster of a domain's problems, the teacher writes one instruction with the background knowledge the cluster depends on, a procedure for that kind of problem, and the common mistakes made on it. This reusable corpus needs only to be built once per domain and requires no run-time teacher access. At inference, a frozen small model retrieves the instructions nearest its question and follows them, with no fine-tuning. Across medicine, law, and mathematics benchmarks, we demonstrate the corpus improves over zero-shot on every task and over few-shot prompting, self-consistency, and other retrieved text on medicine and law. On MedQA the corpus raises mean accuracy by 10.6 points, where retrieved textbook passages raise it by 3.9 and few-shot examples from the same teacher lower it. An error analysis shows that only the background knowledge fixes the questions a small model always gets wrong, and that the procedure and common mistakes fix only the questions where it wavers between options. Our results show that automatically retrieved inference-time procedural guidance and domain knowledge can yield substantial gains for small models.
Figures & tables
Figure 1: One board-exam question with the text a small model reads under passage retrieval and under instruction retrieval. The passages are textbook pages on ulcerative colitis; the fact the question needs (highlighted in both panels) is one sentence among many. The instruction states that fact as background knowledge, orders the steps in the procedure (a plain x-ray first, CT only if the diagnosis stays unclear), and names the common mistakes, ordering CT or colonoscopy first. On MedQA, instructions raise mean accuracy by 10.6 points and passages by 3.9 (Table 1 ). The full instruction is in Appendix A.2 .
Figure 2: Zero-shot and instruction-retrieval accuracy for each student on the three test splits. Each row is one student; the black dot is its zero-shot accuracy, the green dot its accuracy with the corpus, the label the gain in points, and the dashed line GPT-5 zero-shot, the frontier reference. No student loses accuracy on any task, and every student gains on MedQA and MathQA, most on MedQA, where the two 2B students gain 16 and 12 points.
Retrieved text
Student
Zero-shot
Few-shot
Self-consist.
RICP
Passages
Random instr.
Instructions
MedQA
Qwen 3.5 2B
51.8
51.1
57.7
53.8
58.5
51.9
67.6
Qwen 3.5 4B
78.7
76.0
80.1
77.9
79.3
76.7
84.0
Gemma 4 E2B
57.2
55.1
58.1
57.4
62.0
53.2
69.4
Gemma 4 E4B
71.2
69.6
72.6
70.6
73.6
67.9
80.3
Table 1: Test accuracy (%) of each student under each condition. Bold marks the corpus, the background colour the student family, and the Mean row averages the five students. The last four conditions place retrieved text in the prompt; passages exist for MedQA only; random instructions are five drawn at random from the same corpus; self-consistency votes over three samples where every other condition generates once. On MedQA the corpus raises the mean by 10.6 points, more than any other condition, while random instructions leave it below zero-shot on every task. Intervals per cell are in Appendix C .
Figure 3: Questions by how many of five sampled answers are correct, for each student and task. Each bar is one student and its segments, from the bottom, are always wrong, rarely right, usually right, and always right. The always-wrong share depends on the task more than on the student, about 6% on MathQA for every student and between 21% (Qwen 3.5 4B) and 32% (Gemma 4 E2B) on MMLU Law.
Figure 4: Change in the percent of wrong development questions each intervention fixes, against three fresh samples with nothing added, by kind of wrong answer and task, averaged over the four Qwen and Gemma students; whiskers are the range across the four students. A question is fixed when three fresh samples are all correct. Thinking on and self-consistency fix usually-right and rarely-right questions and almost no always-wrong question. The background knowledge Kj is the only part that reaches the always-wrong questions, and which part fixes most differs by task.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Task
k
Size 1
Size 3
Size 8
Size 27
MedQA
1
+5.7
+4.0
+5.5
+0.8
MedQA
3
+9.5
+8.8
+6.6
+4.0
MedQA
5
+11.5
+11.0
+9.1
–
MMLU Law
1
+5.3
− 0.6
+1.2
+0.3
MMLU Law
3
+6.5
+0.6
− 0.3
+0.0
MMLU Law
5
+5.3
+4.7
+5.6
–
Appendix
Table 2: Gain over zero-shot in accuracy points (mean of the four Qwen and Gemma students) on the selection half of the development split, by cluster size (questions per cluster) and number of instructions retrieved k . The chosen cell is size 3 at k=5 ; size 27 was not run at k=5 .
Student
Checkpoint
Context
Generation cap
Thinking cap
Qwen 3.5 2B
Qwen/Qwen3.5-2B
32,768
10,000
16,384
Qwen 3.5 4B
Qwen/Qwen3.5-4B
32,768
10,000
16,384
Gemma 4 E2B
google/gemma-4-E2B-it
16,384
8,192
12,288
Gemma 4 E4B
google/gemma-4-E4B-it
16,384
8,192
12,288
Ministral 3 3B
mistralai/Ministral-3-3B-Instruct-2512
32,768
10,000
–
Appendix
Table 3: Checkpoints, context windows, and generation and thinking caps (tokens) of every student; Ministral 3 3B has no thinking mode. Every sampled run uses temperature 0.7 and top-p 0.95.
Condition
Qwen 3.5 2B
Qwen 3.5 4B
Gemma 4 E2B
Gemma 4 E4B
Ministral 3 3B
Mean of five
MedQA
Zero-shot
51.8
78.7
57.2
71.2
62.5
64.3
Few-shot
51.1 [ − 3.7, + 2.3]
76.0 [ − 4.9, − 0.5]
55.1 [ − 4.5, + 0.5]
69.6 [ − 3.8, + 0.5]
61.3 [ − 3.7, + 1.4]
62.6 ( − 1.6)
Self-consist.
57.7 [ + 3.1, + 8.8]
80.1 [ − 0.6, + 3.5]
58.1 [ − 0.9, + 2.7]
72.6 [ − 0.2, + 3.1]
64.7 [ − 0.4, + 4.8]
66.6 ( + 2.4)
RICP
53.8 [ − 1.1, + 5.2]
77.9 [ − 3.0, + 1.3]
57.4 [ − 2.4, + 2.7]
70.6 [ − 2.7, + 1.5]
59.3 [ − 6.0, − 0.3]
63.8 ( − 0.4)
Passages
58.5 [ + 3.6, + 10.0]
79.3 [ − 1.8, + 2.8]
62.0 [ + 2.1, + 7.5]
73.6 [ + 0.2, + 4.9]
67.2 [ + 2.2, + 7.5]
68.1 ( + 3.9)
Appendix
Table 4: Test accuracy (%) per student under every condition, with the 95% bootstrap interval of the paired difference from zero-shot beneath each value; the last column is the mean of the five students with its difference from zero-shot. Instructions + self-consistency is the vote over the instruction run’s three samples.
Student
Zero-shot
Few-shot
RICP
Passages
Random instructions
Instructions
MedQA
Qwen 3.5 2B
3
2
11
16
10
18
Qwen 3.5 4B
1
0
0
0
2
0
Gemma 4 E2B
1
0
8
2
1
6
Gemma 4 E4B
1
0
0
2
0
1
Ministral 3 3B
4
10
16
14
11
6
Appendix
Table 5: Format failures per cell: greedy replies from which no answer could be read, scored wrong. Passages exist for MedQA only.
Student
n
Always right
Usually right
Rarely right
Always wrong
MedQA
Qwen 3.5 2B
1,272
20.9
30.8
31.0
17.3
Qwen 3.5 4B
1,272
59.4
21.8
10.9
7.9
Gemma 4 E2B
1,272
39.3
18.0
17.5
25.2
Gemma 4 E4B
1,272
57.9
16.0
11.2
14.8
Ministral 3 3B
1,272
34.1
27.1
22.5
16.3
Appendix
Table 6: Percent of development questions per category: always right (five of five samples correct), usually right (three or four), rarely right (one or two), always wrong (none).
Student
Category
Self-consist.
Thinking on
Background
Procedure
Mistakes
Full instr.
MedQA
Qwen 3.5 2B
Usually (392)
70.4
65.3
59.2
45.7
44.4
61.2
Qwen 3.5 2B
Rarely (394)
24.1
25.1
30.5
17.8
17.5
35.5
Qwen 3.5 2B
Always wrong (220)
10.0
10.5
18.6
5.9
11.8
21.4
Qwen 3.5 4B
Usually (277)
78.3
84.8
71.8
60.3
62.1
75.1
Qwen 3.5 4B
Rarely (139)
38.8
51.1
47.5
34.5
40.3
51.8
Appendix
Table 7: Percent of wrong development questions each intervention fixes, per student and category (with the category’s size); the columns are the six interventions of Figure 4 . Intervals per cell are in the released results.
Fact held by
Qwen 3.5 2B
Qwen 3.5 4B
Gemma 4 E2B
Gemma 4 E4B
Ministral 3 3B
Mean of five (P / I)
Wrong zero-shot: % put right
Both texts
274; 57.3 / 67.2
94; 66.0 / 69.1
231; 51.1 / 59.7
130; 59.2 / 62.3
199; 55.8 / 64.8
57.9 / 64.6
Instructions only
160; 28.1 / 54.4
80; 28.7 / 68.8
150; 14.7 / 54.0
104; 19.2 / 65.4
134; 23.9 / 55.2
22.9 / 59.6
Passages only
57; 45.6 / 40.4
32; 50.0 / 21.9
53; 43.4 / 32.1
40; 45.0 / 22.5
43; 37.2 / 34.9
44.2 / 30.3
Neither
123; 23.6 / 22.8
65; 32.3 / 24.6
111; 17.1 / 20.7
93; 17.2 / 23.7
102; 20.6 / 17.6
22.2 / 21.9
Right zero-shot: % put wrong
Appendix
Table 8: Questions grouped by which text holds the fact, split by the student’s greedy zero-shot answer. Each cell is the group’s size, then the percent the five passages (P) put right / the five instructions (I) put right; for questions right zero-shot, the percent each put wrong.
Task
n -gram
Test questions
Instructions
Max shared
MedQA
13
0 of 1,273
0
0
MedQA
8
3 of 1,273
9
3
MMLU Law
13
0 of 540
0
0
MMLU Law
8
2 of 540
2
2
MathQA
13
0 of 2,975
0
0
MathQA
8
13 of 2,975
7
3
Appendix
Table 9: Shared 13-grams and 8-grams between the instructions and the test questions: the test questions flagged, the instructions flagged, and the largest number of 8-grams one test question shares with one instruction.
Task
Shared phrase
Questions
Instructions
MedQA
dopaminergic neurons in the substantia nigra pars compacta
1
2
a non lactose fermenting oxidase positive gram negative
1
1
non lactose fermenting oxidase positive gram negative rod
1
1
holosystolic murmur at the left lower sternal border
1
5
MMLU Law
Appendix
Table 10: Every phrase of eight or more words that a test question shares with an instruction, with the number of test questions and of instructions it appears in.
Student
Options alone
Passages
Instructions
With the question
MedQA
Qwen 3.5 2B
25.6
34.2 ( + 8.6 [ + 5.4, + 11.9])
39.2 ( + 13.6 [ + 10.1, + 16.9])
67.6
Qwen 3.5 4B
26.3
42.9 ( + 16.6 [ + 13.3, + 19.7])
42.5 ( + 16.2 [ + 13.0, + 19.2])
84.0
Gemma 4 E2B
19.6
40.9 ( + 21.3 [ + 18.3, + 24.3])
37.6 ( + 18.0 [ + 14.8, + 21.1])
69.4
Gemma 4 E4B
16.2
41.6 ( + 25.5 [ + 22.5, + 28.4])
30.6 ( + 14.4 [ + 11.5, + 17.0])
80.3
Ministral 3 3B
25.7
39.6 ( + 13.9 [ + 10.6, + 17.2])
42.3 ( + 16.7 [ + 13.5, + 20.0])
73.1
Appendix
Table 11: The question-withheld check, greedy on test: percent of questions on which the student chooses the keyed option from the options alone, from the five retrieved passages (MedQA only), and from the five retrieved instructions, under the same ask, with the question for scale. A reply naming no option counts as not choosing the key. In parentheses, the paired difference from the options alone with its 95% interval.