Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.
Figures & tables
Figure 1: Overview of HeadEdit . Paired completion differences are used to identify a low-rank behavioral subspace. HeadEdit then rescales the projection onto this fixed subspace, producing a state-dependent, vocabulary-wide logit correction while keeping model weights frozen.
Model
Method
Tool Overuse
Over-Refusal
Factual Sycophancy
Over. ↓
Acc. ↑
Safe ↑
Acc. ↑
Anti. ↑
Acc. ↑
Qwen3-4B
Baseline
97.8
57.0
57.3
60.6
37.8
37.8
RePE ( Zou et al., 2023 )
97.8
57.0
59.5
63.4
43.3
43.9
CAA ( Rimsky et al., 2024 )
97.7
57.0
70.2
69.7
45.7
45.7
SADI-Head ( Wang et al., 2025 )
98.9
56.7
68.1
69.1
42.1
40.2
Spherical ( You et al., 2026 )
95.4
58.2
67.0
70.8
37.8
37.2
Table 1: The overall results comparison across three models and three tasks. Over. means tool overuse rate, Safe means safe answer rate, and Anti. means anti-sycophancy rate. All values are percentages. We highlight best results and second-best results, and use the same color for ties.
Figure 2: Results on Qwen3-14B model. Values show the change from the unmodified base model, and higher values are better. HeadEdit improves both measures on all three tasks.
Figure 3: Behavioral information is broadly decodable, but effective control is orientation-specific. Panels (a–c) compare held-out behavioral decodability from the paired-completion subspace and equally ranked random orthonormal subspaces on Qwen3-4B. Shading denotes one standard deviation across ten random subspaces; the dashed line is a linear probe on the full hidden state. Panel (d) evaluates actual steering after matching the RMS magnitude of the centered full-vocabulary logit perturbation.
Figure 4: (a) Test set AUROC of linear probes fit on the training split, using the full prompt-final state, paired rank-8 coordinates V⊤h , or matched random subspaces to predict tool necessity and whether a baseline error is repaired by HeadEdit. Random results show mean ± standard deviation over ten seeds. (b) HeadEdit repair rates among baseline errors, grouped by whether the full-state and V⊤h probes correctly predict the normative tool-use label.
Figure 5: Individual components induce distributed vocabulary-level effects. Bars show the mean prompt-conditioned logit correction induced by a selected component for tool overuse on Gemma3-4B (left) and factual sycophancy on Qwen3-4B (right). Tokens are ordered by decreasing absolute correction magnitude. Muted orange and teal identify readily interpretable task-related tokens whose logits are respectively increased and suppressed; gray tokens are left unannotated rather than assumed to be task-irrelevant. More detailed analyses are provided in Appendix F .
Model
Method
Tool Overuse
Over-Refusal
Factual Sycophancy
Over. ↓
Keep ↑
Acc. ↑
Safe ↑
Unsafe ↓
Acc. ↑
Anti. ↑
Acc. ↑
Qwen3-4B
Baseline
97.8
99.7
57.0
57.3
29.5
60.6
37.8
37.8
LoRA-DPO
92.8
99.6
59.2
67.9
25.0
69.7
54.3
53.0
+ HeadEdit
43.1
73.8
66.4
71.0
38.6
68.6
59.8
59.1
Gemma3-4B-IT
Baseline
32.2
42.2
52.0
63.4
27.3
65.7
25.9
25.3
LoRA-DPO
23.2
40.0
54.0
84.7
34.1
80.0
32.7
31.5
Table 2: Gradient-based alignment results (%) across three models and three tasks. HeadEdit is applied to LoRA-DPO using the settings selected on the corresponding base model, without further tuning. The unmodified model is shown for reference, and bold marks the higher accuracy between LoRA-DPO and LoRA-DPO + HeadEdit.
Figure 6: General ability of Qwen3-4B after steering. Each method directly uses the settings reported in Table 1 , with no further tuning. The dashed line shows the performance of unmodified base model. Error bars in (b) show the range over five seeds. HeadEdit stays close to the base model on both benchmarks, while the other methods show larger drops in some settings.
Figure 7: Sample efficiency of HeadEdit on Qwen3-4B over-refusal. Error bars show the empirical min/max range across five seeds.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Model
qμ(4)
qμ(8)
Qwen3-4B
0.511
0.621
Gemma-3-4B-IT
0.670
0.796
Llama-3.2-3B-Instruct
0.374
0.430
Appendix
Table 3: Mean-direction recovery of centered PCA on When2Tool.
Figure 8: Sensitivity of HeadEdit to subspace rank and intervention gain across three models and three tasks. Rows correspond to models, and columns correspond to behavioral tasks. Each line shows the accuracy change from the unedited model for one rank as the intervention gain α varies. The horizontal dashed line marks no change from the unedited model, and positive values indicate improved accuracy.
Method
Qwen3-4B
Gemma-3-4B-IT
Llama-3.2-3B-IT
Baseline
28.32±0.02
23.86±0.03
47.00±0.02
CAA
28.16±0.03
23.54±0.01
46.79±0.03
RePE
28.25±0.03
23.67±0.05
46.84±0.09
SADI-Head
27.88±0.02
23.37±0.01
45.92±0.01
Spherical Steering
27.75±0.01
23.24±0.03
44.83±0.02
HeadEdit (Ours)
28.32±0.03
23.85±0.01
47.02±0.06
Appendix
Table 5: Steady-state generation throughput (tokens/s; higher is better) on When2Tool. All methods use BF16 inference, batch size one, greedy decoding, and exactly 128 new tokens per prompt. Values are mean ± standard deviation over three runs. LoRA-DPO adapters remain unmerged.
Method
Qwen3-4B
Gemma3-4B-IT
Llama3.2-3B-IT
Baseline
99.7
42.2
98.3
CAA
99.6
41.5
91.8
RePE
99.7
41.5
98.8
SADI-Head
100.0
43.6
11.5
Spherical
100.0
41.3
72.0
HeadEdit
80.9
38.7
86.6
Appendix
Table 6: Necessary tool use rates and unsafe prompt refusal rates across models and intervention methods.
Figure 9: General-capability performance of HeadEdit on Gemma-3-4B-IT and Llama-3.2-3B-IT. Each group uses the HeadEdit configuration selected for the named behavioral task in Table 1 . Gray bars show the unedited model, and blue bars show HeadEdit . Error bars on GPQA-Diamond show the performance range across five seeds.
Qwen3-4B
Gemma-3-4B-IT
Evaluation
Cosine similarity ↑
Explained variance ( R2 ) ↑
Cosine similarity ↑
Explained variance ( R2 ) ↑
Full vocabulary
0.790
0.570
0.685
0.434
Fisher-weighted vocabulary
0.939
0.757
0.743
0.669
Baseline top-10 vocabulary
0.954
0.881
0.886
0.710
DPO-repaired prompts
0.786
0.492
0.687
0.439
Appendix
Table 7: Predicting DPO logit updates from base-model behavioral coordinates. A rank-8 ridge predictor is evaluated on held-out prompts using different vocabulary supports and on the subset of prompts repaired by DPO.
Figure 10: DPO-update predictability persists across edit magnitudes. Held-out prompts are partitioned from the smallest (Q1) to largest (Q4) DPO-induced logit changes. The same emulator is evaluated without refitting. Its explanatory power is not concentrated on near-zero updates, while the Q4 decline exposes variation not captured by V⊤h .
Figure 11: Tail perplexity of Base and HeadEdit generations across four models and three behavioral tasks. Perplexity is evaluated by the unedited model over tokens following the first generation position.
Figure 12: HeadEdit and Baseline Output Comparison on Qwen3-4B in Over-refusal Dataset
Figure 13: HeadEdit and Baseline Output Comparison on Gemma3-4b in Factual Sycophancy Dataset
Figure 14: HeadEdit and Baseline Output Comparison on Gemma3-4b in Tool Overuse Dataset
Figure 15: HeadEdit and Baseline Output Comparison on Qwen3-4b in Factual Sycophancy Dataset
Figure 16: HeadEdit and baseline output comparison on Qwen3-14B for over-refusal.
Figure 17: HeadEdit and Baseline Output Comparison on Qwen3-14b in Tool Overuse Dataset
Figure 18: HeadEdit and Baseline Output Comparison on Llama3.2-3b in Over-refusal Dataset