Injecting skills into a frozen language model currently costs a million parameters and a reinforcement-learning pipeline. We introduce DecSteer, a System-1 decision operator trained by behavior cloning that lowers this cost by roughly two orders of magnitude. The default operator uses 330K parameters to match a 1.33M-parameter operator trained with reinforcement learning, exceeds or achieve comparable performance, while collapsing 3,685-token deliberation into a 6-token decision with no loss in accuracy. A rank-4 variant with 23K parameters, 1/58 of the strongest published skill operator, suffices for SearchQA and near-suffices for LiveMath, where higher rank still helps; the same recipe transfers across five tasks and three backbones, with out-of-distribution gains persisting on LiveMath problems released months after training. The gap to prior work is trainability, and it is set jointly by initialization and architecture. The initialization of prior operators zeroes the gradient of both large factor matrices at the first optimization step, whereas our zero-initialized output projection inside a shared low-rank backbone receives a gradient immediately, which a gradient-flow probe confirms directly. The gain isn't chain-of-thought compression. 23 of 57 LiveMath points beat the base model's best-of-8 sampling, and a logit-lens probe shows the operator amplifies the answer along the model's existing late-layer pathway, not writing it earlier. Gains track the base model's headroom across 13 base-task pairs, and skills compose as approximately linear operators that can be added, interpolated, and hot-swapped at inference time.
Figures & tables
Figure 1: A decision does not need a million parameters. Left: LiveMath accuracy vs. trainable parameters. A 23K-parameter model trained by behavior cloning matches a 1.33M-parameter one trained with RL, while the same BC fails on the prior one. Right: It collapses median output length from 3,685 to a 6-token decision.
Figure 2: Overview. a: A shared low-rank backbone is injected at D=8 depths with per-depth input-dependent gate g ; the operator is identity at initialization. b: The gate can be phase-conditioned on prompt vs. generation via zero-init increments. c: The identity anchor steers optimization: with up=0 , the output projection gets gradient from step one, while baselines V=0 keeps both large matrices frozen until V moves. d: Trained skills act as composable residual objects.
LiveMath (Acc)
SearchQA (EM)
CSQA (Acc)
OpenBookQA (Acc)
DocVQA (ANLS)
Qwen3.5-4B
Base
0
22.4
68.1
80.7
90.6
86.9
Text Skill
0
23.4
67.4
78.5
89.0
88.4
SkillOpt
–
52.0
71.2
–
–
89.0
Soft Skill
–
64.5
76.4
–
–
88.2
LoRA (BC)
1.38M
19.4
69.1
–
–
93.7
Table 1: Main results. Ours is a low-rank residual operator trained by behavior cloning on answer tokens only. KV-Skill, LoRA, Text Skill, SkillOpt, and Soft Skill rows quote published numbers; the Base rows are their published base scores. Rank rows are trained from scratch with the same recipe. In the smaller-backbone blocks, where no published baselines exist, the KV-Skill row is our faithful reimplementation trained with the identical BC recipe (Appendix A ). Ours entries are means over 3 seeds; a uniform operator strength of 1.0 yields the same conclusions.
Figure 3: Left: accuracy versus trainable parameters per task. DecSteer (BC) at 0.33M reaches or exceeds every KV-Skill variant at 1.33M, while KV-Skill’s own BC control (same signal, same examples) stays at the base. Right: median generated tokens, base versus DecSteer. LiveMath drops from 3,685 to 6 tokens (36 × wall-clock), DocVQA from 289 to 11 (10.3 × ), with accuracy unchanged or higher; SearchQA already answers directly.
Figure 4: Left: the rank–parameter–score ladder on SearchQA: 7.7K parameters already exceed the published 1.33M RL result, and the operator size scales linearly with the backbone width, so smaller models take smaller skills. Right: out-of-distribution generalization. A bundle trained only on problems released through February 2026, evaluated zero-shot on new LiveMath problems from April–June 2026 (zero id overlap): 79.0 versus base 25.4, and 7.5 points above the official released KV-Skill RL artifact under the identical evaluation protocol.
Figure 5: Left: gradient-flow probe over the first optimization steps. With up=0 , the output projection B has gradient Θ(ds) at step 1; with KV-Skill’s V=0 anchoring, both large matrices have exactly zero gradient while only the small middle matrix moves. Right: accuracy versus rank, trained from scratch with BC. The knee sits between rank 2 and 4 on LiveMath; SearchQA saturates at rank 1. Rank-4 (23K parameters) reaches the 1.33M level.
Figure 6: Left: decomposing the LiveMath gain. Base greedy 23.4; forced direct answering via prompt 38.7; base best-of-8 sampling 57.3; DecSteer 80.7. The last step exceeds the base model’s sampling budget: on 43 problems the operator is correct while all 8 base samples are wrong. Right: logit-lens readout of answer readability by layer. The answer becomes readable only within the last 1–6 layers in both arms; the operator amplifies the existing late-layer signal (3.4–6.4 × , more for weaker bases) rather than writing the answer earlier.
Figure 7: Left: the headroom trend, an empirical relation. Skill gain versus base headroom ( 1− base score) over 13 base–task pairs on three backbones; the relation is monotone, and saturated bases (CSQA, OpenBookQA on 4B) sit at zero as the trend expects. Right: skill algebra. Adding, interpolating, and reordering two skills preserves the anchor skill in all 18 cells (zero interference); the one forbidden zone is a negative coefficient on a mode-switching skill, which reverses the modality it switches.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Check
Result
Identity anchor ( θ0 , up =0 )
Output bit-identical to base (max ∣Δ∣≤0.0625 , bf16 rounding)
Identity at scale =0 (n=400 / n=50)
Item-identical to the no-bundle base
Train ∩ test overlap (LiveMath)
0 duplicated ids
LiveMath v7 vs. train split
0 / 319 id overlap (2026-04/05/06)
Gold label distribution (LiveMath)
≈ uniform; majority class 28.2%
Prediction distribution (LiveMath)
≈ uniform; no single-label collapse
Appendix
Table 2: Protocol and integrity checks behind the reported numbers: exact identity anchors for fair comparison, leakage and degeneracy screens, budget audits, and the import validation of the official KV-Skill artifact used as a baseline.
Figure 8: Operator strength sweeps (mean ± SD over seeds). Stars mark each task’s best strength.
Task
Method
Seeds
Score ± SD
KV-Skill
SearchQA
Ours (BC)
3
80.9 ± 0.6
80.1
SearchQA
Ours (GRPO scratch)
3
80.7 ± 0.9
80.1
SearchQA
Ours (GRPO from skill)
1
80.6
80.1
LiveMath
Ours (BC)
4
79.6 ± 0.8
79.0 ± 3.9
LiveMath
Ours (GRPO from skill)
1
80.7
79.0 ± 3.9
LiveMath
Ours (GRPO scratch)
1
71.8
75.6 ± 4.1
Appendix
Table 3: Behavior cloning matches GRPO under KV-Skill’s own protocol, at 1/4 the parameters. All BC seeds meet or exceed the KV-Skill RL reference (80.1 / 79.0). On LiveMath, GRPO from scratch trails both its KV-Skill counterpart and the BC warm start: the starting point matters more than the algorithm.
Model
Task
Base
Ours (3 seeds)
KV-Skill (3 seeds)
Δ
Llama-3.2-3B
SearchQA
57.1
77.6 ± 1.3
76.5 ± 0.1
+1.1
Llama-3.2-3B
LiveMath
16.9
71.2 ± 2.0
66.1 ± 1.4
+5.1
Llama-3.2-3B
CSQA
71.3
72.7 ± 0.9
71.7 ± 0.6
+1.0
Llama-3.2-3B
OpenBookQA
74.6
78.2 ± 0.7
78.9 ± 0.8
−0.7
Qwen3.5-0.8B
SearchQA
36.1
72.9 ± 0.8
70.1 ± 0.7
+2.8
Qwen3.5-0.8B
LiveMath
–
70.7 ± 3.3
62.6 ± 3.1
+8.1
Appendix
Table 4: Cross-carrier comparison: the KV-Skill operator is reimplemented exactly ( V=0 identity init) and both operators are trained and evaluated under identical data, prompts, and decoding. Ours wins 6 of 8 cells and is never worse, with 1/4 the parameters.
Task
Signal
Examples
Wall-clock
Peak mem.
LiveMath (4B)
BC
35
12 min
13.9 GiB
SearchQA (4B)
BC
400
33 min
–
DocVQA (4B)
BC
107
29 min
27.9 GiB
CSQA (4B)
RFT
400
–
–
SearchQA (4B)
GRPO
300 steps
45 min
12.6 GiB
LiveMath (4B)
GRPO
300 steps
77 min
13.6 GiB
Appendix
Table 5: Training cost on a single A100-40GB. The BC recipe used for all main results trains in minutes because there are no rollouts; GRPO rows are the protocol-matched controls of Table 3 . Inference additionally accelerates generation by up to 36 × (LiveMath), 10.3 × wall-clock (DocVQA).
Figure 9: SearchQA calibration before and after skill injection: ECE (left) and answer-ranking AUROC (right) across the three backbones.
Figure 10: Left: multimodal extension on DocVQA. The identical BC recipe transfers to document VQA: 4B gains +2.3 ANLS (92.5 to 94.8, higher in 3/3 seeds) and 0.8B gains +23.8 (57.9 to 81.8), with a second confirmation of mild overshoot at full strength. Right: behavior on saturated 4B bases. Paired flip counts are almost perfectly symmetric (53/52, 58/59, 82/80, 7/7) and all skill arms sit within about a point of the base, the headroom trend’s expectation for a saturated base.
Figure 11: Left: RFT control on CSQA. Even when the training target contains full chains of thought (330/400 sampled solutions kept by the verifier, median target 1,445 characters), the operator still converges to 6-token direct answers with accuracy unchanged: the direct-answer mode is the operator’s least-effort solution, not an artifact of short targets. Right: RL stability. The published entropy bonus overwhelms the KL anchor by roughly 100 × once entropy rises, driving divergence (reward 0.86 to 0.08); removing it restores stability. This supports using BC, which needs none of these controls.
Figure 12: Commitment-point probe: applying a suffix intervention from generation step k measures how early the answer is fixed. The height of the curve at k=0 tracks the eventual skill gain across the four tasks probed and follows the headroom trend, so it serves as an auxiliary screening diagnostic; the informative quantity is the curve’s height, not its onset.