COMPASS: Finding Where Reasoning Lives in Language Models
Organizations: University of California, Riverside, CA 92521. Work performed during a summer internship at Lawrence Livermore National Laboratory · Lawrence Livermore National Laboratory, Livermore, CA 94550.
Abstract
Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model's own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.
Figures & tables
| Llama-3.1-8B | Qwen3-4B | Gemma4-12B | |||||||
| GSM8K | SVAMP | MATH-500 | GSM8K | SVAMP | MATH-500 | GSM8K | SVAMP | MATH-500 | |
| Standard Prompt | 66.8 (93) | 81.0 (48) | 42.4 (345) | 76.2 (68) | 81.0 (22) | 55.2 (250) | 49.4 (67) | 74.7 (3) | 68.4 (347) |
| ITI Li et al. (2023) | 73.8 (132) | 68.7 (45) | 23.6 (409) | 71.9 (45) | 44.3 (78) | 30.6 (56) | 22.3 (52) | 64 (46) | 30.3 (872) |
| Fractional Reasoning Liu et al. (2025) | 79.9 (214) | 81.0 (119) | 42.2 (524) | 88.3 (220) | 89.3 (118) | 77.8 (561) | 65.4 (55) | 85.6 (7) | 72.7 (371) |
| COMPASS (ours) | 83.7 (143) | 85.7 (81) | 45.2 (515) | 89.8 (138) | 90.7 (68) | 77.8 (404) | 88.0 (140) | 79.3 (51) | 79.8 (366) |
| CoT (oracle) | 83.4 (205) | 85.0 (157) | 47.0 (788) | 92.3 (303) | 91.7 (224) | 83.2 (924) | 96.1 (338) | 94.3 (261) | 91.4 (730) |
| Llama-3.1-8B | Qwen3-4B | |||
|---|---|---|---|---|
| GSM-Plus | HARP | GSM-Plus | HARP | |
| Standard Prompt | 60.4 (136) | 21.5 (366) | 65.6 (82) | 35.4 (396) |
| COMPASS (ours) | 71.1 (180) | 22.7 (750) | 80.4 (164) | 49.1 (575) |
| CoT (oracle) | 72.7 (247) | 27.1 (1020) | 84.1 (360) | 56.9 (1505) |
| GPT-5.4-mini (oracle) | 88.4 (385) | 69.7 (1509) | 88.4 (385) | 69.7 (1509) |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Description |
|---|---|
| GSM8K [ 14 ] | Grade-school word problems with integer answers; 7,473 train and 1,319 test problems. |
| SVAMP [ 15 ] | 1,000 one-step arithmetic word problems built by perturbing existing problems so that surface cues do not determine the operation; 700/300 train/test split of the public release (ChilleD/SVAMP). Train and test share no question text. |
| MATH-500 [ 34 ] | The 500-problem subset of the MATH test set [ 2 ] , seven subjects (prealgebra to precalculus) at five difficulty levels, LaTeX answers. No fitting is done on these 500 problems: the Llama selection is fitted on the 7,500-problem MATH train split, and the Qwen3 selection on the remaining 4,500 MATH test problems (MATH test minus MATH-500, matched on normalized problem text and verified disjoint). |
| GSM-Plus [ 16 ] | 10,552 adversarial perturbations of GSM8K test problems across eight perturbation types. The 1,319 critical-thinking rows are unanswerable by construction (gold “None”) and are dropped, leaving 9,233 evaluated problems. |
| HARP [ 17 ] | 4,780 short-answer problems from US national mathematics competitions (AMC, AIME, and others), six difficulty levels, LaTeX answers. No train split by design; steered with the MATH selection. |