Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models
Organizations: Skolkovo Institute of Science and Technology · AI Foundation and Algorithm Lab · Moscow Institute of Physics and Technology · Artificial Intelligence Research Institute (AIRI) · CNRS, Universit´e Paris Cit´e
Abstract
Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer internally. We expose this latent knowledge via the Query--Key (QK) score, defined for an attention head as the inner product between the last-token query and the key at the end-of-line token following option , evaluated before rotary positional embedding is applied. Its argmax identifies a universal class of select-and-copy heads in middle layers that perform option selection through semantic query--key alignment, mechanistically distinct from induction and copy-suppression heads (Olsson et al., 2022): they are invariant to label symbols, and solve a synthetic task with zero surface overlap---properties no positional-copy account explains and that critically require stripping RoPE. Across 24 models from 1.5B to 72B parameters (LLaMA-2/3/3.1/3.3, Qwen-2.5, Gemma, Phi-3.5, DeepSeek-R1-Distill), a single head's QK-score exceeds the model's own zero-shot accuracy by up to pp on HellaSwag and pp on HaluDialogue; causal zero-ablation collapses MCQA accuracy to near-random. To remove any dependence on labeled validation data, we introduce an unsupervised HeadScore that ranks heads from unlabeled inputs and recovers the supervised top- heads on every tested model. Against four positional-debiasing baselines (e.g., PriDe, Wiegrefe, Wang), QK-score is complementary by construction: debiasing re-weights output logits, whereas QK-score reads the model's selection from a middle-layer head before decoding. We release a one-line drop-in HeadScore script and per-model head indices, making every result one-command reproducible across all 24 models and four benchmarks.
Figures & tables
| Language | Max Acc | Min Acc |
|---|---|---|
| English | 1.000 | 0.995 |
| Italian | 1.000 | 0.995 |
| French | 1.000 | 0.985 |
| Russian | 1.000 | 0.990 |
| LLaMA… | LLaMA… (chat, instruct) | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | 3-8B | 3-70B | 3.1-8B | 3.1-70B | 3-8B | 3-70B | 3.1-8B | 3.1-70B | |
| MMLU | |||||||||
| Baseline | Acc | 60.3 | 75.3 | 59.1 | 72.9 | 60.5 | 78.2 | 60.4 | 73.8 |
| PA | 50.4 | 68.8 | 48.6 | 65.7 | 47.7 | 70.1 | 50.1 | 67.8 | |
| QK-score | Acc | 61.0 | 74.5 | 62.5 | 75.9 | 63.0 | 77.9 | 64.9 | 79.4 |
| PA | 51.5 | 66.0 | 50.7 | 68.5 | 49.3 | 67.9 | 55.2 | 74.4 | |
Appendix figures & tables54 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Method | MMLU | Cosmos QA | Hellaswag QA | Halu Dialogue | |
|---|---|---|---|---|---|---|
| Baseline | Acc | 59.5 0.6 | 74.3 1.2 | 58.1 2.2 | 35.9 2.1 | |
| Qwen2.5 | PA | 49.3 0.4 | 67.8 1.6 | 48.6 2.9 | 25.7 1.7 | |
| -1.5B | QK-score | Acc | 57.0 1.3 | 72.1 0.9 | 58.7 1.0 | 38.3 0.9 |
| PA | 45.9 1.1 | 64.2 0.8 | 48.6 0.3 | 27.5 0.8 | ||
| Baseline | Acc | 29.2 1.5 | 37.3 1.2 | 29.4 0.2 | 26.5 1.1 | |
| LLaMA2 | PA | 10.6 0.5 | 17.3 0.9 | 9.1 0.1 | 7.2 1.2 |
| Model | Method | MMLU | Cosmos QA | Hellaswag QA | Halu Dialogue | |
|---|---|---|---|---|---|---|
| Baseline | Acc | 56.9 2.9 | 71.4 2.7 | 56.7 3.7 | 38.4 2.5 | |
| Qwen2.5 | PA | 45.9 3.2 | 64.2 3.2 | 47.3 3.8 | 27.2 2.5 | |
| -1.5B | QK-score | Acc | 55.3 2.0 | 70.6 1.3 | 56.1 1.8 | 38.2 2.5 |
| PA | 43.3 1.9 | 61.3 1.4 | 45.8 2.9 | 23.6 5.7 | ||
| Baseline | Acc | 1.8 | 29.6 5.0 | 27.3 1.3 | 23.9 1.2 | |
| LLaMA2 | PA | 10.0 1.5 | 11.5 4.2 | 8.1 0.9 | 5.7 0.8 |
| …-shot prompting | |||||||
|---|---|---|---|---|---|---|---|
| Method | 0 | 1 | 2 | 3 | 4 | 5 | |
| MMLU | |||||||
| Baseline | Acc | 59.0 | 63.3 | 63.3 | 62.6 | 62.7 | 63.7 |
| PA | 48.4 | 52.6 | 52.9 | 52.1 | 52.0 | 53.3 | |
| PRIDE | Acc | 62.4 | 64.0 | 63.3 | 63.0 | 63.4 | 64.1 |
| PA | 52.9 | 53.8 | 54.2 | 53.7 | 53.5 | 54.1 | |
| …-shot prompting | |||||||
|---|---|---|---|---|---|---|---|
| Method | 0 | 1 | 2 | 3 | 4 | 5 | |
| MMLU | |||||||
| Baseline | Acc | 26.7 | 39.1 | 43.1 | 43.7 | 44.1 | 43.8 |
| PA | 8.9 | 21.3 | 26.2 | 27.4 | 28.5 | 28.4 | |
| PRIDE | Acc | 15.5 | 36.9 | 39.8 | 40.8 | 41.5 | 42.7 |
| PA | 5.7 | 20.8 | 24.2 | 24.6 | 25.6 | 28.9 | |
| Dataset | Best (Layer, Head) |
|---|---|
| MMLU | (16, 19) , (17, 24) , (16, 26), (17, 26) |
| HaluDialogue | (14, 5), (14, 21), (14, 2), (17, 24) |
| HellaSwag | (14, 5) (17, 24) , (17, 29), (16, 19) |
| CosmosQA | (17, 24) , (16,19) , (16, 26), (17,29) |
| Dataset | Best (Layer, Head) |
|---|---|
| MMLU | (14, 24) , (15, 4), (17, 0), (14, 20) , |
| HaluDialogue | (14, 29), (14, 24) , (14, 26) |
| HellaSwag | (15, 5), (15, 4), (18, 10), (14, 20) |
| CosmosQA | (14, 24) , (15, 5), (15, 4), (15, 23), (14, 20) |
| MMLU | CosmosQA | HaluDialogue | HellaSwag | ||||||
| SFT | LLaMA | SFT | LLaMA | SFT | LLaMA | SFT | LLaMA | ||
| 0-shot | Baseline | 0.493 | 0.267 | 0.863 | 0.311 | 0.923 | 0.211 | 0.352 | 0.265 |
| QK | 0.488 | 0.336 | 0.840 | 0.414 | 0.905 | 0.371 | 0.393 | 0.330 | |
| 1-shot | Baseline | 0.476 | 0.391 | 0.836 | 0.393 | 0.474 | 0.309 | 0.716 | 0.288 |
| QK | 0.472 | 0.407 | 0.810 | 0.5 | 0.858 | 0.366 | 0.800 | 0.371 | |
| 2-shot | Baseline | 0.477 | 0.431 | 0.643 | 0.591 | 0.823 | 0.342 | 0.795 | 0.306 |
| MMLU | CosmosQA | HaluDialogue | HellaSwag | |||||
|---|---|---|---|---|---|---|---|---|
| Cloze | QK | Cloze | QK | Cloze | QK | Cloze | QK | |
| 0-shot | 0.38 | 0.35 | 0.49 | 0.46 | 0.42 | 0.40 | 0.52 | 0.38 |
| 1-shot | 0.40 | 0.39 | 0.51 | 0.50 | 0.45 | 0.42 | 0.52 | 0.35 |
| 2-shot | 0.39 | 0.40 | 0.48 | 0.51 | 0.46 | 0.42 | 0.53 | 0.38 |
| 3-shot | 0.39 | 0.40 | 0.48 | 0.57 | 0.45 | 0.37 | 0.53 | 0.43 |
| 4-shot | 0.39 | 0.39 | 0.53 | 0.54 | 0.46 | 0.39 | 0.53 | 0.43 |
| Dataset | both | QK / baseline | QK / baseline | both |
|---|---|---|---|---|
| MMLU | 61.09 | 7.99 | 6.07 | 24.85 |
| CosmosQA | 79.26 | 8.64 | 2.64 | 9.46 |
| HellaSwag | 76.69 | 5.91 | 5.26 | 12.14 |
| HaluDialogue | 48.06 | 15.07 | 6.44 | 30.43 |