cs.AIOct 5, 2026

COMPASS: Finding Where Reasoning Lives in Language Models

Authors: Pratyay Dutta, Kowshik Thopalli, Vivek Narayanaswamy

Organizations: University of California, Riverside, CA 92521. Work performed during a summer internship at Lawrence Livermore National Laboratory · Lawrence Livermore National Laboratory, Livermore, CA 94550.

Abstract

Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model's own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Math Reasoning in LLMs is Organized by Approach, Not Topic

    Sep 22, 2026Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri +1Mathematical Reasoning BenchmarksLarge Language Model Tool Use

  2. Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

    Aug 4, 2026Denys Pushkin, Albert Q. Jiang, Aryo Lotfi +2Chain-of-Thought ReasoningInstruction-Tuned Models

  3. Manifold-Guided Attention Steering

    May 20, 2026Ian Li, Kapilesh Guruprasad, Raunak Sengupta +3Linear Activation SteeringLLM Reasoning Strategies