Large language models (LLMs) have been rapidly improving in long-context tasks, powered by Chain-of-Thought (CoT) reasoning. However, the internal mechanisms underlying this improvement remain unclear. We investigate these mechanisms through a needle-in-a-haystack (NIAH) counting task, where an LLM is asked to count the number of records dispersed in a long text. Across twelve model comparison groups, Thinking (or reasoning) improves counting accuracy over Non-thinking, with pronounced gains at larger counts. This motivates our mechanistic analysis, which identifies two contrasting mechanisms: (i) broad retrieval, where Non-thinking models broadly attend to multiple needles; (ii) targeted retrieval, where Thinking models use enumeration in CoT traces to successively retrieve needles. Targeted retrieval concentrates attention on individual needles and is accompanied by more compact internal representations. Moreover, causal intervention analysis suggests that Thinking models use the CoT trace to maintain and update an internal counter as needles are successively retrieved, even without explicit numbering. In small controlled experiments, both retrieval mechanisms and counter states emerge under standard autoregressive training. Together, our results connect long-context retrieval with representation geometry of counting, supporting a state-tracking account of CoT reasoning.
Figures & tables
Qwen3
Gemma-4
Nemotron
GLM †
Ministral †
L
Mode
4B
8B
14B
32B
E4B
12B
26B-A4B
31B
v2-9B
3-4B
4/Z1-9B
3-8B
1k
Non-thinking
43.3
60
50
96.7
36.7
73.3
100
100
40
46.7
20
73.3
Thinking
100
100
100
100
100
100
100
100
100
96.7
100
96.7
20k
Non-thinking
3.3
30
30
46.7
33.3
10
43.3
63.3
3.3
13.3
13.3
16.7
Thinking
90
96.7
100
100
43.3
83.3
93.3
100
33.3
73.3
46.7
70
Table 1 : Thinking improves counting accuracy across models. Table shows parsed exact accuracy (%), N=5 , over 30 paired seeds. † Separate checkpoints.
Figure 1 : Counting accuracy across counts and context lengths. A,B. Median and interquartile range across 12 model groups; lines are smoothed guides. C,D. Accuracy over 30 paired seeds: Non-thinking (rose, solid) and Thinking (blue, dashed), with darker colors for larger N . Lines connect measured values. Qwen is evaluated with YaRN disabled.
Figure 2 : Broad retrieval in Non-thinking mode. A. Needle-attention shares for each model’s six highest-ranked heads, normalized within each N=10 prompt and averaged over 20 seeds. B. Broad retrieval scores Bh across counts 2–10; highlighted points mark the frozen Qwen Top-32 and Gemma Top-6 sets. Gemma includes global-attention layers only.
Figure 3 : The form–retrieve–consolidate pathway. A. Schematic. B. Clean-to-corrupted restoration and reverse patching of needle spans. C. Retrieval-head ablation vs. layer-matched random and original controls. D. Proportion of predictions matching the source count after answer state patching.
Figure 4 : Targeted retrieval in Thinking. A. Needle-normalized attention for one head per model ( N=10 ); circles mark needle k+1 at Marker k . B. Mean Th across seeds; highlights mark Qwen Top-128 and Gemma Top-6. Gemma includes global-attention layers only.
Figure 5 : The retrieve–update–read pathway. A. Schematic. B. Retrieval-head ablation vs. layer-matched random and original controls. C. Counter state patching from a later trace item to an earlier one, using traces without indices (Qwen L19; prompted Gemma L21). Later columns report continuation success conditional on success at the preceding step and a remaining item. D. Proportion of predictions matching the source count after answer state patching.
Figure 6 : Learning dynamics of two synthetic counting models. A. Free-running count accuracy. B,C. Layer-wise broad retrieval scores in Non-thinking and targeted retrieval scores in Thinking. The dashed line marks the switch to task-output loss at step 1,500. See Appendix A for evaluation details.
Appendix figures & tables48 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Relative noise between adjacent counts. Non-thinking answer-query states at 10k tokens: Qwen3-8B L25 and Gemma-4-E4B L36, selected by cross-validated NCC (Appendix D.2 ). Relative noise is pooled within-count standard deviation divided by projected mean separation (Eq. ( 1 )); larger values indicate poorer separation. Directions use 20 fitting seeds; estimates use ten held-out seeds per count. Lines connect observations. Shading shows pointwise 95% intervals from 10,000 bootstrap draws, independently resampling whole-seed profiles in each split, refitting directions, and fixing layers.
Figure 8 : Count-model fits to unsmoothed cross-model medians. All 14 counts and eight lengths are shown. Curves use Eqs. ( 2 ) and ( 3 ).
Mean accuracy (%)
Non-thinking
Thinking
Comparison group
NT
T
CV R2
RMSE
CV R2
RMSE
Qwen3-4B
31.8
94.4
0.750
17.9
0.782
4.9
Qwen3-8B
45.3
96.5
0.739
17.0
0.778
2.8
Qwen3-14B
35.5
98.6
0.663
20.7
0.293
2.3
Qwen3-32B
47.7
98.0
0.816
16.3
0.508
2.5
Gemma-4-E4B
29.2
79.4
0.730
18.2
0.927
7.1
Appendix
Table 2 : Observed means and held-count fits. The 1k–20k and 1k–100k grids contain 3,360 and 7,140 requests per mode per group, respectively. RMSE: percentage points. † Separate checkpoints.
Figure 9 : Complete 12-group comparison at 1k–20k. Every panel includes 14 counts, eight lengths, and both modes. Points are 30-request accuracies; solid Gaussian curves and circles denote Non-thinking, dashed step-success curves and triangles Thinking. Colors identify length. Fits are separate for each group. † Separate checkpoints.
Figure 10 : Qwen3-32B: all counts and lengths, YaRN disabled. Each panel fixes N and shows all 17 lengths from 1k to 100k. Rose circles/solid lines are Non-thinking; blue triangles/dashed lines are Thinking. Points are 30-seed means, shading shows pointwise 95% Wilson intervals, and lines connect measurements without smoothing. Truncations count as errors.
Figure 11 : Gemma-4-31B: all counts and lengths. The display matches Fig. 10 . Gray strips separate the 1k–20k and 25k–100k evaluation batches; lines and intervals do not cross that gap. All 14 counts and both modes are retained, including local accuracy reversals.
Figure 12 : Qwen3-32B long-context accuracy with and without YaRN. Panels show Non-thinking (rose) and Thinking (blue). Solid circles: YaRN disabled; dashed triangles: YaRN configured. Each point pools 14 counts with 30 seeds each (420 requests); truncations and unparseable outputs count as errors. Lines connect observations without smoothing. The runs also differ in execution settings (request batch size 2 versus 1), so gaps do not isolate YaRN’s causal effect.
Figure 13 : Candidate length dependence of effective step failure. Dots are the free-by-length estimates from Eq. ( 3 ); curves are the five joint fits in Eq. ( 4 ). All fits use accuracy observations, not the dots. Panel A uses 1k–20k medians; B,C use each model’s complete 1k–100k data. Vertical scales differ.
Thinking data
Linear γ
Log γ
Linear λ
Log λ
Reciprocal q
Across-model median
0.86/4.9
0.53/9.0
0.86/5.0
0.53/9.0
0.85/5.0
Qwen3-4B
0.74/5.4
0.43/7.9
0.73/5.4
0.43/7.9
0.73/5.4
Qwen3-8B
0.71/3.1
0.41/4.5
0.71/3.2
0.41/4.5
0.71/3.2
Qwen3-14B
0.36/2.2
0.24/2.4
0.36/2.2
0.24/2.4
0.36/2.2
Qwen3-32B
0.61/2.3
0.37/2.9
0.61/2.3
0.37/2.9
0.61/2.3
Gemma-4-E4B
0.85/10.1
0.52/18.2
0.84/10.3
0.51/18.3
0.84/10.6
Appendix
Table 3 : Length-model validation: R2 /RMSE (percentage points). The first 13 rows leave out each length in the 1k–20k grid; the next two do so over 1k–100k. The final two are frozen short-to-long predictions. All five candidates have three parameters and retain all 14 counts. Higher R2 and lower RMSE indicate better predictions. † Separate checkpoints.
Analysis
Data and split
Head scores
20 selection seeds, N=2 –10: 180 prompts.
Needle end geometry
N=1 –10: 1,100 fitting endpoints from 200 prompts and 550 held-out endpoints from 100 prompts.
Count classification
20 selection seeds, five-fold seed-grouped cross-validation. Both readouts use N=1 –10: 1,100 running-index endpoints and 200 answer-query states.
Prompt-span steering
Ten initially correct N=3 prompts; all layers.
Cue/domain controls
Cue: ten paired seeds 1234–1243 at N=10 . Domain: fit on 20 city fitting seeds; test on ten held-out seeds, N=1 –10, in each domain.
Head ablation
20 seeds 1316–1335, N=1 –5: 100 prompts.
Appendix
Table 4 : Data and splits for Non-thinking experiments, per model unless specified. Eligibility and invalid-output rules accompany each analysis.
Figure 14 : Complete broad-retrieval scores. A. All 1,152 Qwen heads. B. All 56 Gemma global-attention heads. Cells show mean Bh on 180 selection prompts. Horizontal axes show heads; vertical axes show layers from shallow to deep. Both panels use a linear 0–1 scale. Outlines mark frozen Qwen Top-32/Gemma Top-6 membership.
Figure 15 : Count readout across layers. A. Running index at needle endpoints. B. Final count at the answer query. Both cover N=1 –10 and use five-fold seed-grouped cross-validation on 20 selection seeds: 1,100 states in A and 200 in B. Accuracy averages recall over the ten labels. Solid/dashed curves denote nearest-centroid/logistic classifiers; dotted lines mark 10% chance. Labeled markers identify the NCC-selected layers, with logistic accuracy and then earlier layers breaking ties.
Figure 16 : Cue and domain controls. A,B: 100 needle end states per condition at N=10 in the first three components of a jointly fitted unwhitened PCA6 (Qwen L16/Gemma L14). C,D: 100 answer-query states per domain (city, flower, animal) in a PCA3 estimated from city examples. Small points are states; connected large markers are class means. Circles/triangles denote cue present/absent in A,B; circles/triangles/squares denote city/flower/animal in C,D. Percentages give fitting-population variance explained. Conditions share the projection and display scale within each panel.
Model
City
Flower
Animal
Qwen L24
67%
50%
56%
Gemma L39
58%
43%
53%
Appendix
Table 5 : Frozen city-trained count readout. Nearest-centroid accuracy on 100 held-out prompts per topic ( N=1 –10; ten seeds). Chance is 10%.
Figure 17 : Retrieval-head ablation. A: Qwen; B: Gemma. Ranked heads (solid circles) versus layer-matched random banks (dashed diamonds). Effect is the absolute change in generated count divided by the true count N . Bands show seed-bootstrap intervals.
Figure 18 : Answer state patching controls. A: Qwen; B: Gemma. Proportion of predictions matching the source count after answer state patching. Different-count (solid), same-count (dashed), and self (dotted) patches use fixed correct-input pairs across layers. Controls are scored against the corresponding different-count source. Shading gives pointwise 95% seed-bootstrap intervals.
Figure 19 : Answer state transfer and removal. A: proportion of predictions matching the source count. B: normalized excess error from aligned removal. Seed-bootstrap intervals (50,000 draws in B).
Figure 20 : Serial interventions across the three stages. A: restoration and retrieval/answer removal effects. B: interaction and remaining repair. All contrasts use absolute expected-count error divided by N before averaging; both panels share the same percentage scale. Bars show seed-bootstrap intervals.
100 held-out traces per topic; projections and probes fit on the city fitting set.
Head ablation
Ten common final transitions with explicit ranks in both models, one per held-out seed; shared across head-count tests. Qwen Top-128/Gemma Top-6.
No-index readout / state transfer
20 fitting / ten held-out traces per model, N=10 ; natural Qwen and prompted Gemma. Readout uses 200/100 states; held-out patches use k=4,6,8 , both directions, and three scopes.
Answer state patching
40 common directed pairs from 39 prompts; four per held-out seed, with both endpoints correct in both models.
Appendix
Table 6 : Data and splits for Thinking experiments, per model unless specified. Repeated conditions share seed clusters.
Figure 21 : Complete targeted-retrieval scores. A. All 1,152 Qwen heads (15/20 eligible selection seeds). B. The 56 Gemma global-attention heads (19/20). Scores use the model-specific formats and queries above. Horizontal axes show heads; vertical axes show layers from shallow to deep. Both panels use a linear 0–1 scale, shared with Enumeration. Outlines mark frozen Top-128/Top-6 membership.
Figure 22 : Count readout across layers. A. Running index at trace-item endpoints. B. Final count at the answer query. Both analyses cover N=1 –10. Curves pool predictions from five seed-grouped cross-validation folds (20 seeds/model): 1,107 Qwen/1,045 Gemma running states and 200 answer states each. Solid/dashed curves denote NCC/logistic classifiers; labeled markers identify NCC-selected layers. Blue denotes Qwen and orange Gemma; dotted lines mark 10% chance.
Figure 23 : Counter-state and answer state PCA in the full Thinking cohort. A,B. Qwen/Gemma running index states. C,D. Final count states. Small points show held-out states, with opacity indicating viewing depth; connected large markers show class means. Color denotes k or N . Axes use PC1–3 fitted using the selection data; percentages report their combined explained variance.
Model / endpoint
Layer
City
Flower
Animal
Qwen / running k
L19
544; 58.9/68.1
546; 58.6/61.7
528; 55.9/70.1
Gemma / running k
L17
519; 71.8/78.0
507; 70.7/78.8
524; 72.4/72.5
Qwen / final N
L26
100; 100/100
100; 98.0/98.0
100; 100/100
Gemma / final N
L37
100; 70.0/71.0
100; 70.0/69.0
100; 77.0/75.0
Appendix
Table 7 : Count readout across needle domains. Cells report state count; NCC/logistic balanced accuracy (%). Each domain contains 100 held-out trajectories per model. City-only fitting and layer selection precede transfer; chance is 10%.
Figure 24 : Count geometry across needle domains. A,B. Qwen/Gemma running states. C,D. Final count states. Circles/triangles/squares denote city/flower/animal; small points show held-out states and connected large markers show class means. Colors encode k or N . All three domains share the projection fitted on city data within each panel; percentages give fitting-population variance explained by PC1–3.
Figure 25 : Head-count dependence in the frozen retrieval banks. A,B. Qwen next-item/final count effects. C,D. Gemma effects (ten transitions/model). Solid circles ablate frozen selected prefixes; dashed diamonds average three layer-matched random banks excluding the selected bank. Axes show ablated-head count and failure increases from clean continuation; shading gives pointwise 95% seed-bootstrap intervals. Masks persist through free generation.
Figure 26 : Counter-state scope controls by direction. A,B. Forward transfer. C,D. Backward transfer. Columns show Qwen L19 and Gemma L21. Each scope uses the same 30 pairs per panel (ten seeds, k=4,6,8 ). Filled/hollow points show source/self-patch successor adoption; bars give pointwise 95% seed-bootstrap intervals.
Figure 27 : Trace dependence of final count readout. A. Qwen. B. Gemma. Exact-count accuracy for Original, prompt-record blanking, and whole-trace blanking (100 unfiltered prompts/model). Error bars give pointwise 95% seed-bootstrap intervals.
Analysis
Population and split
Full-cohort geometry
200 fitting / 100 held-out inputs; each model retains its available parsed running states and all answer states.
Head scores / ablation
20 selection seeds; ten model-common final transitions, one per held-out seed.
40 model-common directed pairs; four pairs per seed, full layer sweeps.
Appendix
Table 8 : Data and splits for Enumeration experiments, per model unless specified. Repeated conditions share seed clusters.
Figure 28 : Complete targeted-retrieval scores. A. All 1,152 Qwen heads. B. The 56 Gemma global-attention heads. Scores average within seed, then across 20 selection seeds/model. Horizontal axes show heads; vertical axes show layers from shallow to deep. Both panels use a linear 0–1 scale, shared with Thinking. Outlines mark frozen Top-128/Top-6 membership.
Figure 29 : Count readout across layers. A. Running index at trace-item endpoints. B. Final count at the answer query. Curves pool predictions from five seed-grouped cross-validation folds (20 seeds/model): 1,102 Qwen/1,067 Gemma running states and 200 answer states each. Solid/dashed curves denote NCC/logistic classifiers; labeled markers identify NCC-selected layers. Blue denotes Qwen and orange Gemma; dotted lines mark 10% chance.
Figure 30 : Counter-state and answer state PCA in Bullet enumeration. A,B. All 550 Qwen/528 Gemma held-out running index states. C,D. All 100 final count states/model, including incorrect original answers. Small points show held-out states, with opacity indicating viewing depth; connected large markers show class means. Color denotes k or N . Axes use PC1–3 fitted using the selection data; percentages report their combined explained variance.
Figure 31 : Head-count dependence in the frozen retrieval banks. A,B. Qwen next-item/final count effects. C,D. Gemma effects (ten transitions/model). Solid circles ablate frozen selected prefixes; dashed diamonds average three random banks. Axes show ablated-head count and failure increases from clean continuation; shading gives pointwise 95% seed-bootstrap intervals. Masks persist through free generation. All doses and eight truncated formal generations remain (four/model); random banks match layers except at Qwen Top-128.
Figure 32 : Counter-state scope controls and continuation. Columns show Qwen L19 and Gemma L21. A,B. Filled/hollow points show source/self-patch successor adoption at each scope (30 trials/direction; ten seeds, k=4,6,8 ). C,D. Item-span next-step success conditional on the entire preceding source-implied city prefix being correct and another item remaining. Steps 1–2 include k=4,6,8 ; steps 3–4 include k=4,6 . Circles/solid lines denote forward transfer; diamonds/dashed lines denote backward transfer. Bars and shading give pointwise 95% seed-bootstrap intervals. Failures and truncations remain in the step for which they are eligible.
Figure 33 : Answer state transfer and trace dependence. A. Proportion of predictions matching the source count after full-state (solid) or self (dashed) patches at the single answer-query token, 40 directed pairs/model at every layer. B. Exact-count accuracy for Original, prompt-record blanking, and whole-trace blanking (100 unfiltered inputs/model). Blue denotes Qwen and orange Gemma. Shading and error bars give pointwise 95% seed-bootstrap intervals.
A. One occurrence ( N=1 )
Query
<BOS> <CountChar> S c D <Sep>
Shuffled input excerpt
… ,r c ␣G …
Target positions
232
Counts by character
S : 0, c : 1, D : 0
Thinking output
<Think> <Sep> c </Think> <Ans> <1> <EOS>
Non-thinking output
<Ans> <1> <EOS>
Appendix
Table 9 : Examples of synthetic inputs. Three actual test-region inputs and their expected outputs in both modes. The input characters are randomly permuted, so these excerpts do not preserve words or sentences. Within each example, both modes receive the same query prefix and 256-character input. Excerpts retain every target occurrence and nearby non-target characters; ellipses omit only non-target characters. Positions are one-based. Boldface marks target occurrences, ␣ denotes an input space, and \n denotes an input newline.
Setting
Value
Model and data
Transformer depth / heads per layer
4 / 8
Hidden / head dimensions
512 / 64 (approximately 12.7M parameters)
Feed-forward network
Width 2,048; GELU with tanh approximation
Normalization / dropout
Pre-LayerNorm / 0 throughout the model
Position encoding / capacity
RoPE, base 10,000 / 384 token positions
Appendix
Table 10 : Synthetic training and evaluation settings. Both modes use the same settings unless the output format or loss component is specified. There is one training run per mode. Learning-rate values refer to one-based optimizer steps.
Figure 34 : Training and counting behavior. A/B: full-sequence, token-weighted cross-entropy on evaluation inputs generated from the training (solid) and validation (dashed) regions of the original corpus. Validation-region inputs are excluded from parameter updates. The dotted line marks the objective change at step 1,500; this evaluation metric remains fixed. C: final-checkpoint answer accuracy, with ten evaluation inputs per count and mode. Missing numeric answers are incorrect.
Figure 35 : Retrieval-head scores. A: Non-thinking broad score Bh at the final-answer query. B: Thinking targeted score Th at trace retrieval queries. Each cell shows one head’s score at step 10,000, averaged equally over 100 teacher-forced evaluation inputs. Rows are layers L1–L4; columns are heads H1–H8. The panels use separate color scales, matching Fig. 41 in Appendix G.4 , with no normalization across heads.
Figure 36 : Retrieval head ablation. A: final count accuracy on all 100 Non-thinking inputs, for Top-1–4 broad heads. B: next-needle accuracy on 88 eligible Thinking inputs, for Top-1–8 targeted heads. Lines show selected sets, control means, and clean baselines; shading spans the control sets. The axes each span 40 percentage points; B is truncated to 60–100%.
Figure 37 : Count geometry. Cross-validated NCC selects L2/L4/L4/L2 for A/B/C/D. Separate PCA projections show 550 running index states (A/B) and 100 answer states (C/D) per mode. Outlined class means are connected in label order. The three components explain 50.9/48.1/74.0/71.7% of selection variance in A/B/C/D. All panels use the same camera view and equal coordinate units within each panel.
Figure 38 : Count decoding across layers. Balanced NCC accuracy for running index (A) and final count (B), using teacher-forced states. Squares mark layers selected by grouped cross-validation. Preprocessing and centroids are fitted only on selection inputs.
Figure 39 : Sources of the next needle. A: prompt and trace schematic. B: next-needle accuracy after prompt or trace ablation, with controls at non-target prompt positions. All 100 inputs are retained, including 30 without prior trace items.
Figure 40 : Trace continuation and final-answer interventions. A: input-averaged source-continuation exact match for an L1 source-item transplant (circles), self patches/clean runs (squares), and norm-matched orthogonal controls (triangles). Conditions use the same eight inputs and 26/30/26/20 eligible pairs for lengths 1/2/3/4. Error bars on source transplants show 95% percentile intervals from 10,000 input-bootstrap resamples; control curves show means. B: Thinking answer accuracy after prompt or trace ablation on 100 generated prefixes. Controls ablate equally many non-target prompt positions; the dotted line marks clean accuracy. C: agreement with the source run’s clean prediction on 180 pairs using teacher-forced prefixes.
Figure 41 : Attention and count decoding during training. All panels use a linear training-step axis and the same 100 evaluation inputs. A–D: attention scores for all 32 heads across 101 checkpoints; rows retain head identity and white lines separate layers. A/B share the broad-score scale; C/D use a fixed 0–1 scale for targeted and successor-like attention. E/F: balanced NCC accuracy for running index and final count at eleven checkpoints. Layers are selected by final-checkpoint cross-validated NCC and remain fixed; transforms and centroids are fitted on selection states at each checkpoint. Vertical dotted lines mark the objective switch at step 1,500; horizontal dotted lines in E/F mark 10% chance accuracy.
Figure 42 : Final-answer accuracy by position and category count. A: requested record position (30 inputs per point). B: true target-category count, pooling city and flower questions (60 per point). Blue denotes Qwen and orange Gemma; solid circles denote Thinking and dashed squares Non-thinking. Aggregate 95% CIs appear in Table 11 .
Task
Model
Non-thinking
Thinking
Kth record
Qwen
13.00 [11.33, 15.00]
97.67 [94.67, 100.00]
Kth record
Gemma
26.33 [23.00, 29.67]
78.00 [70.00, 85.67]
Category count
Qwen
32.33 [28.33, 36.33]
96.00 [93.67, 98.33]
Category count
Gemma
30.33 [27.00, 34.00]
61.67 [54.67, 68.67]
Appendix
Table 11 : Natural final-answer accuracy (%). Each condition includes all 300 inputs from 30 seeds. Brackets give 95% confidence intervals (CIs).
Figure 43 : Head scores for both additional tasks. Task groups pair broad ( Bh , Non-thinking) and targeted ( Th , Thinking) scores; rows show Qwen (blue) and Gemma (orange). Cells represent heads, with one-based layer and head indices. Each panel has an independent, unclipped linear scale from zero to its labeled maximum; compare numerical scores using the colorbars.
Figure 44 : Top- K head ablation. Absolute accuracies for Selected (solid circles) and Random (dashed open squares; mean of three banks). A–D: Thinking targeted ablation, measuring next-entity accuracy. E–H: Non-thinking broad ablation, measuring final-answer accuracy. Columns show Qwen (blue) and Gemma (orange); rows separate tasks. Shading shows pointwise 95% CIs. Qwen’s broad axes are logarithmic. Clean accuracies for Qwen/Gemma are 100.00/97.00% for kth-record retrieval and 94.78/69.36% for category counting in the targeted assay; the corresponding broad-assay accuracies are 13.00/26.00% and 25.00/29.00%.
Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete patterns. We propose the Hidden CoT Detection Score (HCDS), a comparative behavioral and mechanistic signal measuring whether neutral-prompt behavior aligns more closely with explicit CoT or explicit no- CoT. Here, hidden CoT operationally denotes this neutral-prompt CoT-like alignment; HCDS does not directly observe or prove an unexposed reasoning trace. On GSM8K, HCDS is significantly positive for both Qwen3-4B variants (Thinking +1.87, p=1.2×10−7; Instruct +1.41, p=1.9×10−4), replicates across a different inference stack and quantization within 0.08 (+1.80 and +1.45), and is not significantly positive in seven of eight length-adjusted calibration-control cells. The unadjusted score produces large positive scores on single-step arithmetic and numeric factual lookup. The variants also respond differently to no-CoT instructions: Instruct complies from the prompt alone, whereas Thinking continues reasoning and requires intervention. These findings show stronger, less prompt-conditional CoT-like behavior in the reasoning-tuned model, consistent with but not proof of latent reasoning. HCDS thus investigates latent reasoning without relying on models' self-reported traces.
Armaan Singh, Ryan Trinh Le, Jasmine Kaur +5
1Stanford University · 2York University · 3Facebook
Chain-of-Thought (CoT) prompting often improves the reasoning performance of large language models (LLMs), but the internal signal that triggers this behavior remains poorly understood. Leveraging the sparse features captured by Sparse Autoencoders (SAEs), we propose a systematic framework to analyze and intervene on the internal representations of LLMs, identifying a small set of latent features that are linked to reasoning behavior and can be causally tested through targeted intervention. Across multiple model families and reasoning benchmarks, we show that steering one or a small number of reasoning-related latent features can substantially induce reasoning behavior without explicit CoT prompting, achieving accuracy comparable to CoT. We further show that the identified features are not tied to particular wording patterns or verbosity, and confirm their causal role in reasoning through suppression experiments that impair performance even under CoT prompting. These results suggest that CoT prompting activates specific latent features to trigger reasoning, and that targeted intervention on these features offers an alternative pathway to elicit efficient reasoning behavior without explicit CoT prompting. Code is available at https://github.com/Zhenghao-He/LatentCoT.
Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often produces long, redundant traces. Recent training-free early-exit methods shorten these traces by choosing an intermediate point to stop reasoning. We study one such strategy that injects an end-of-think token (EoT, </think>) at this point to trigger the reasoning-to-answering transition, and find that the injected EoT does not always induce a clean answering phase. Answering-phase generation can continue before the model regenerates another EoT, with the span preceding this regenerated EoT scaling with the reasoning tokens saved by early exit and exhibiting continued reasoning behavior. We call this spurious CoT termination, where reasoning-like generation continues into the answering phase. We hypothesize that insufficient attention to the injected EoT contributes to spurious CoT termination and probe this hypothesis with Exit-token Attention Biasing (EAB). Across four LRMs, five benchmarks, and two early-exit methods, increasing attention to the injected EoT reduces spurious CoT termination and answering-phase length. These results reveal a limitation of controlling LRMs by externally matching their explicit think-block format. Inserting the EoT token conforms to this format but does not by itself guarantee the intended reasoning-to-answering transition. Our code is available at https://github.com/Seunghee-Koh/Spurious-CoT-Termination.
Seunghee Koh, Sungjae Choi, Minchan Kwon +2
Korea Advanced Institute of Science and Technology, South Korea