In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at https://github.com/gzxiong/task_operator.
Figures & tables
Figure 1: Overview of the Task Operator framework: extracting per-head affine parameters (w,b) from ICL forward passes ( left ), replaying them via an updated output projection at zero-shot inference ( middle ), and sparsifying to the highest-importance sites ( right ). The yellow blocks denote the context tokens, and the blue blocks denote the tokens of the input query.
Model
Method
Lexical
Algorithmic
Reasoning
Translation
Linguistic
Uppercase
Reverse
Deduplicate
GSM8K
MATH500
GPQA
Qwen3 (4B)
ICL
69.53
73.44
100.00
43.75
24.22
89.39
55.80
37.88
ZSL
0.00
0.00
0.00
0.00
0.00
74.53
23.60
17.68
TV
43.75
49.22
92.19
1.56
0.00
81.05
28.00
21.21
FV
0.00
17.97
94.53
0.00
16.41
73.16
26.00
15.15
ICV
0.00
0.00
0.00
0.00
0.00
74.83
23.80
20.20
Table 1: Test accuracy (%) across models and tasks. ICL denotes performance with K=8 examples provided. The scores without gray shading denote the performance with no in-context demonstrations.
Figure 2: Per-site importance for three tasks on Llama3.1-8B. Brighter cells indicate higher importance. Token labels are color-coded by position type.
Task
Method
m=8
m=16
m=24
m=32
ICL
Uppercase
TV
87.50
88.28
87.50
88.28
100.00
TO
100.00
100.00
100.00
100.00
Reverse
TV
5.47
6.25
6.25
6.25
48.44
TO
56.25
54.69
51.56
51.56
Deduplicate
TV
2.34
2.34
2.34
2.34
37.50
TO
19.53
26.56
28.91
28.12
Table 2: Effect of the number of validation prompts m on test accuracy (Llama3.1-8B, K=8 fixed). ICL ( K=8 ) is shown as a reference baseline.
Figure 3: Effect of the number of demonstrations K on test accuracy (Llama3.1-8B, m=32 ). TO (batch) averages operators from disjoint 8-shot batches.
Setting
Lexical
Algorithmic
Reasoning
Translation
Linguistic
Uppercase
Reverse
Deduplicate
GSM8K
MATH500
GPQA
ICL
69.53
73.44
100.00
43.75
24.22
89.39
55.80
37.88
Template + Content + Output
67.19
71.09
100.00
39.06
17.19
86.20
51.20
33.33
Template + Output
67.19
71.09
100.00
8.59
10.94
85.44
53.60
29.29
Template + Content
66.41
70.31
100.00
17.97
2.34
77.26
50.00
23.74
Template ( w + b )
67.19
71.88
100.00
7.81
9.38
79.61
47.20
22.22
Table 3: Position-type and affine component ablations on Qwen3-4B. Top rows vary position scope. Bottom rows isolate the scaling ( w ) and bias ( b ) components at template positions.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Task
CV(w)<
cos(b)>
0.1
0.2
0.5
0.9
0.8
0.5
Qwen3-4B
Translation
70.4
83.9
95.4
88.1
98.0
100.0
Reverse
77.3
90.0
97.1
97.7
99.6
100.0
GSM8K
71.4
85.1
94.3
83.5
92.7
99.2
Qwen3-8B
Translation
71.9
84.9
96.0
86.2
97.4
100.0
Reverse
76.6
89.5
96.3
97.8
99.7
100.0
Appendix
Table 4: Cross-sample stability across ICL prompts. Each row reports the fraction of (layer, token, head) sites (in %) whose per-prompt scaling factor satisfies CV(w) below the threshold and whose per-prompt bias satisfies mean pairwise cos(b) above the threshold.
Model
Task
Broadcast (%)
Per-rank (%)
Δ (pp)
Qwen3-4B
Translation
67.2
67.2
−0.0
Reverse
39.1
36.7
−2.3
GSM8K
86.2
81.2
−5.0
Qwen3-8B
Translation
67.2
67.2
−0.0
Reverse
60.2
50.0
−10.2
GSM8K
78.6
50.8
−27.8
Appendix
Table 5: Cross-token stability via downstream accuracy. The per-rank variant uses offset-aware (w,b) for the first 12 offsets in content and output; broadcast averages across offsets to yield one (w,b) per category. Δ : per-rank minus broadcast accuracy.
Model
Task
Pearson( w )
Mean cos(b)
Median cos(b)
Qwen3-4B
Translation
0.981
0.844
0.896
Reverse
0.992
0.953
0.967
GSM8K
0.993
0.912
0.947
Qwen3-8B
Translation
0.982
0.835
0.891
Reverse
0.992
0.954
0.968
GSM8K
0.993
0.908
0.945
Appendix
Table 6: Cross-demonstration-set stability. Pearson( w ) is computed once per (model, task) cell over all flattened (layer, token, head) sites. Cosine similarity of b is computed per site, with mean and median aggregated across sites within each cell.
Figure 4: Test accuracy under top- p sparsification on three algorithmic tasks across four models.
Model
Method
Uppercase
Reverse
Deduplicate
Qwen3 (4B)
ICL
100.00
43.75
24.22
TO (Full)
100.00
39.06
17.19
TO (Sparse @ p⋆ )
100.00
39.84
21.09
Sites kept (%)
100
60
60
Qwen3 (8B)
ICL
100.00
67.19
76.56
TO (Full)
100.00
60.16
65.62
Appendix
Table 7: Test accuracy (%) of the dense and val-NLL-sparsified task operator on the three algorithmic tasks. p⋆ is selected by minimizing validation NLL over p∈{0.2,0.4,0.6,0.8,1.0} . Sites kept is the fraction of sites retained at p⋆ .
Model
Method
Translation
Deduplicate
GPQA
t0
t1
t2
t0
t1
t2
t0
t1
t2
Qwen3 (4B)
ICL
69.53
67.97
67.19
24.22
14.84
18.75
37.88
29.80
29.29
ZSL
0.00
0.00
0.00
0.00
0.00
0.00
17.68
16.16
17.68
TV
43.75
41.41
11.72
0.00
0.78
0.78
21.21
30.30
23.23
TO (Ours)
67.19
64.84
62.50
17.19
15.62
9.38
33.33
29.29
28.79
Qwen3 (8B)
ICL
68.75
67.97
65.62
76.56
47.66
35.16
34.85
32.83
34.85
Appendix
Table 8: Test accuracy (%) under three matched prompt templates. t0 denotes the canonical Input / Output format, t1 the Q / A format, and t2 the compact arrow format. TO and TV are extracted and evaluated under the same template. Bold marks the best score among ZSL, TV, and TO.
Model
Method
MATH-500
GPQA-Diamond
Qwen3 (4B)
ICL
55.80
37.88
ZSL
23.60
17.68
TO (in-domain)
51.20
33.33
TO (transfer)
51.20
21.72
Qwen3 (8B)
ICL
61.20
34.85
ZSL
19.00
21.21
Appendix
Table 9: Cross-task transfer of task operators extracted from GSM8K to MATH-500 and GPQA-Diamond. The results compare transferred operators with zero-shot inference, in-domain task operators, and standard in-context learning. Underlined entries indicate methods that improve over the zero-shot baseline.