Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, minimal traces. Answer correctness and trace validity nearly coincide in distribution but decouple out of distribution: on the hardest instances, 31.6% of correct answers have invalid traces, over half of which pass all syntactic and arithmetic checks but fail semantic dependency checks. We then intervene on trace supervision. Non-minimal training traces induce non-minimal outputs, while re-asking the same problem with a different query reveals computations inherited from the original query, weakening minimality as evidence of selective planning. Shuffling tokens in 10% of training trace sentences preserves near-clean accuracy even out of distribution despite no trace passing verification. Swapped training traces likewise retain high in-distribution accuracy. We discuss the implications of these findings for chain-of-thought monitoring and interpretation in the context of AI safety.
Figures & tables
in-distribution
out of distribution
training trace
op ≤15
op =15
op =20
op =21
op =22
op =23
clean
100.0
98.8
83.3
77.8
67.8
58.0
swapped
81.7
19.9
16.5
16.9
16.8
17.4
swapped, op-matched
81.4
19.2
16.1
16.4
16.7
17.1
Table 1: Accuracies for iGSM with clean, swapped traces, and swapped traces with operations matched. The accuracies for each operation in-distribution is in Appendix Table 6
in-distribution
out of distribution
training trace
op ≤15
op =15
op =20
op =21
op =22
op =23
clean
100.0
98.8
83.3
77.8
67.8
58.0
10% shuffled
99.9
98.5
83.9
78.5
69.6
60.1
30% shuffled
98.6
74.9
43.5
37.3
31.2
27.8
50% shuffled
93.5
42.9
18.4
16.8
14.6
15.0
75% shuffled
83.1
0.6
0.0
0.0
0.0
0.0
Table 2: Table showing results for training model with shuffled prefix sentences of the trace. The results are compared with models trained on clean traces, on swapped traces, and on no trace (the model emits only the answer).
op =15
op =20
op =23
correct answers
4047
3410
2376
with an invalid trace
20
256
750
share of correct answers with invalid traces (%)
0.5
7.5
31.6
Syntactic/arithmetic failures (total)
1
56
366
share of invalid traces (%)
5.0
21.9
48.8
trace does not parse
1
54
352
Table 3: Invalid traces accompanying correct answers, grouped by the checks they fail. Table 9 in the appendix extends it to all five levels, splits the row “a step uses the wrong inputs” by how the answer survived, and points to an example of each row.
slack in the problem
emitted trace
training traces
op
unneeded params
none (%)
extra params / trace
minimal (%)
minimal (clean)
15
2.66
37.8
0.00±0.07
99.9
20
1.84
49.9
0.01±0.45
98.7
23
1.35
60.0
0.13±2.02
93.8
90% non-minimal
15
2.66
37.8
1.20±2.47
73.9
20
1.84
49.9
0.99±2.09
75.2
Table 4: Minimality of emitted traces under minimal and non-minimal training supervision. Slack is the number of stated parameters the query does not need (mean over 4,096 problems, and the share of problems with none). We report the unnecessary parameters per trace (mean ± sd) and the share of traces with no unnecessary parameters.
world
op
query
accuracy
irrelevant
padded (%)
unnec. params / trace
standard
10
original
99.9
3
0.0
0.00
standard
20
original
84.5
1
0.4
0.00
standard
15
re-asked
89.6
2–9
2.6
0.14
standard
20
re-asked
86.9
2–9
3.7
0.18
wide
10
original
100.0
11
0.4
0.01
wide
20
original
67.6
6
1.3
0.01
Table 5: Minimality of the clean model’s traces in a widened world and for re-asked questions, 1,024 problems per cell; op is that of the problem as generated. Irrelevant sentences are stated sentences the query does not need. Padded traces and unnecessary parameters per trace are counted over correctly answered problems; a padded trace defines at least one unnecessary parameter.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
op
problems
clean
swapped
swapped, op-matched
1
553
100.0
100.0
97.8
2
493
100.0
100.0
98.6
3
483
100.0
99.8
99.6
4
419
100.0
100.0
100.0
5
396
100.0
96.5
97.2
6
325
100.0
89.2
91.4
Appendix
Table 6: Accuracies for iGSM with clean, swapped traces, and swapped traces with operations matched, for each in-distribution operation count.
op
problems
clean
10%
30%
50%
75%
100%
swapped
no trace
1
553
100.0
100.0
100.0
99.8
100.0
99.6
100.0
100.0
2
493
100.0
100.0
100.0
99.8
100.0
96.3
100.0
100.0
3
483
100.0
100.0
100.0
98.6
100.0
90.7
99.8
100.0
4
419
100.0
100.0
100.0
98.6
98.8
89.5
100.0
100.0
5
396
100.0
100.0
100.0
97.2
97.0
84.6
96.5
98.2
6
325
100.0
100.0
100.0
96.0
93.8
61.2
89.2
96.9
Appendix
Table 7: In-distribution answer accuracy (%) per number of operations for the clean model, the shuffled-prefix models ( p=10 to 100 ), the swapped model, and the model trained with no trace.
op =15
op =20
op =21
op =22
op =23
correct
wrong
correct
wrong
correct
wrong
correct
wrong
correct
wrong
trace valid
4027
20
3154
132
2788
115
2181
141
1626
109
trace invalid
20
29
256
554
397
796
595
1179
750
1611
Appendix
Table 8: For the clean model, a quadrant of trace validity versus answer correctness for all the operations. We find that at larger operations, the correlation between trace validity and answer correctness drops.
op =15
op =20
op =21
op =22
op =23
example
correct answers
4047
3410
3185
2776
2376
with an invalid trace
20
256
397
595
750
share of correct answers (%)
0.5
7.5
12.5
21.4
31.6
Syntactic/arithmetic failures (total)
1
56
146
267
366
share of invalid traces (%)
5.0
21.9
36.8
44.9
48.8
trace does not parse
1
54
141
259
352
Ex. 3
Appendix
Table 9: Table 3 extended to all five operation levels. The row “share of correct answers” is the number of invalid traces divided by the number of correct answers. The two “share of invalid traces” rows divide each group’s total by the number of invalid traces, the row “with an invalid trace”. The last column points to an example of each row in Appendix D .
clean model
steps
setup
op =15
op =20
op =21
op =22
op =23
Ye et al. (reported)
100k
trace valid too
99.1
91.8
87.9
84.0
76.8
run A
100k
used in all tables
98.8
83.3
77.8
67.8
58.0
trace valid too
98.3
77.0
68.1
53.2
39.7
run B
100k
second run
99.1
81.3
71.7
57.2
47.3
run C
200k
twice the steps
99.7
87.4
80.4
68.5
58.5
trace valid too
99.6
83.9
73.6
57.3
41.3
Appendix
Table 10: Accuracy (%) of our clean models against the values reported by Ye et al. (2025a) for iGSM-med; all runs use their recipe. Ye et al. count a solution only when its trace also passes their checker; the rows marked “trace valid too” apply the same criterion to our models. Runs A and B differ by ten points at op =23 .
in-distribution
out of distribution
training trace
seed
op ≤15
op =15
op =20
op =21
op =22
op =23
swapped
42
81.7
19.9
16.5
16.9
16.8
17.4
43
79.9
19.1
16.4
16.3
16.6
17.2
swapped, op-matched
42
81.4
19.2
16.1
16.4
16.7
17.1
43
81.2
18.3
16.6
16.3
16.6
16.7
30% shuffled
42
98.6
74.9
43.5
37.3
31.2
27.8
Appendix
Table 11: Answer accuracy (%) of a second training seed for the models trained on swapped, shuffled and non-minimal traces. The seed-42 rows are the models reported in the main text; the seed-43 rows use the same recipe. Accuracies agree within five points at every level and the ordering of the models is unchanged.
School of Electrical and Computer Engineering, Georgia Institute of Technology, USA · Department of Electrical and Computer Engineering, Bogazici University, Turkey