Humans can apply ideas learned in one context to substantially different situations. We call this process of searching for and constructing novel recombinations of learned relationships to solve new problems \textit{out-of-distribution (OOD) reasoning}. The capacity for large language models (LLMs) to support OOD reasoning underpins proposals for generally intelligent systems. However, this property is challenging to measure because most problems admit many decompositions, some involving shallow subproblems and others involving subproblems that may have been memorized from the training data. This makes it difficult to determine how much compositional reasoning a model must actually perform. Uncertainty about training distributions, how to measure a datapoint's distance from a training distribution, and the exponential number of ways to decompose most tasks make it intractable to robustly measure any modern LLM's OOD-reasoning capacity on standard natural-language tasks. In this work, we introduce a benchmark in which a model progressively solves a Boolean circuit over GF(2) from data. This benchmark is both minimal, isolating OOD reasoning from these confounding factors, and general, as any computable function can be represented as a sufficiently large GF(2) polynomial. We find that standalone models' next-step accuracy collapses as depth grows. In contrast, tool use through the synthesis and execution of code prevents this collapse in both small and frontier LLMs. These results indicate that synthesizing tools play a crucial role in supporting OOD reasoning.
Figures & tables
Figure 1: The left panel provides an abstract depiction of training. The dots represent training datapoints in an idea space, and the arrows denote relationships linking these datapoints. The center panel zooms in on the region around a few datapoints and visualizes the problem-solving process during inference. The empty circles represent an initial query, and the LLM follows arrows corresponding to relationships learned in-distribution to reach an answer. If both the reasoning path and the answer lie near a training datapoint (dashed concentric circles), the process is considered in-distribution reasoning; otherwise, it is considered out-of-distribution reasoning. The long gray arrow illustrates the synthesis and application of a tool within the reasoning process, which expands the model’s OOD reasoning capabilities.
Figure 2: Evaluating one reasoning step at a given depth. Filter for datapoints decryptable using the supplied prefix Pg : ag+1=1 and aj=0 for j>g+1 . This selection does not require knowing the missing support Sg+1 . The LLM cancels the known contribution ⨁j=1gajMj(v) from the labels and predicts t^ . Exact-match frequency estimates γg , compared with chance 1/(d−1p) . Repeated correct steps could reconstruct the circuit; the evaluation scores individual proposals from supplied prefixes.
Figure 3: Left: Only estimators with access to both history and data sustain reliable next-step prediction. We plot step success γg (the probability mass assigned to the correct next monomial) as a function of depth g for each estimator class. Curves show the mean over 2000 generated circuits at each depth, with shaded Jeffreys intervals. Right: As both g and p increase, estimators with incomplete information tend to collapse toward zero probability mass on the correct next monomial. Only estimator A (history+data) consistently maintains high γg . In contrast, B (data-only) and D (partial) often collapse toward zero, while C (history-only) remains near chance. The heatmap is based on 200 generated circuits for each hyperparameter setting.
Figure 4: Left: Tool use improves next-step prediction for small instruct models, although accuracy still declines with depth. We plot step success γg versus circuit depth g for Qwen3-2507 4B- and 30B-Instruct under adversarial sampling ( p=12 , d=4 ). Center: With tools, 4B-Thinking sustains high accuracy and 30B-Thinking recovers near-perfect performance at larger depths; without tools, both models approach chance. Right: Tool use sustains frontier-model accuracy at large depths. At g=255 , Opus 4.8 and Sol 5.6 achieve near perfect accuracy with tools, versus negligible without tools. Solid and dotted curves denote performance with and without tools, respectively; shading indicates 95% Jeffreys intervals, and the dashed line marks chance ( 1/220 ). Together, these results show that tool use substantially mitigates depth-induced collapse on this benchmark, and doesn’t experience any collapse on the range tested for thinking models.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Additional turns do not restore accuracy after execution is disabled. Left: original tool-enabled accuracy at g=31 . Right: next-step accuracy at g=32 with retained history but no further execution; dashed lines show the original no-tool reference at g=32 . The panels use different vertical scales to make the 0 – 0.4% control results visible. Values are the reported aggregate accuracies; colored regions represent Jeffery’s intervals.
Figure 6: Iterative tool use is essential for small models. For Qwen3-4B-Thinking , next-step accuracy improves as the model is allowed more code execution and revision cycles. A single generated program is often insufficient at large depth, while repeated execution feedback enables the model to repair errors and recover the prefix-conditioned computation.
Figure 7: Token usage is similar with and without tools. The plot shows the number of input, output, and cached tokens used to predict the next term at g=31 . Token counts are averaged over 500 runs and divided by the corresponding average without tools for the same model and depth. Although each tool invocation includes all previously generated tokens as input, the shorter reasoning sequences keep overall token usage similar across conditions, with a maximum deviation of approximately 35% . If accessing through API, tool use is likely still cheaper across the board even when it consumes more tokens because cached tokens are substantially discounted and code execution incurs minimal cost and latency. Reusing cached tokens can also reduce inference time, allowing tool-enabled models to solve these problems faster.
Figure 8: As both g and p increase, the probability of an estimator with imperfect information begins to collapse to zero. Only Estimator A is able to consistently produce the next monomial. The above heatmap was constructed through generating 200 circuits for each combination of hyperparameters and computing the corresponding γg for each p . This diagram shows the performance when not using the adversarial dataset construction.
Figure 9: The figure above shows how γg changes for Estimators A , B , C , and D when the data is biased adversarially (as in the results so far) preventing identification of monomials from frequency statistics, and without de-biasing.
Figure 10: Slopes below 1 without tools imply diverging error with depth; slopes near 1 with tools imply error rate independent of problem depth. Estimator Dk is the partial-information solver from Sec. 3.1 given access to only the first k revealed prefix monomials, with predicted next-step accuracy γD(p,g,k) . For each (model,g) pair, k⋆(g)=argmink∣γobs(g)−γD(p,g,k)∣ is the prefix length under which Dk would reproduce the LLM’s observed accuracy, i.e. the effective prefix the LLM behaves as if it integrated. Shaded bands are 95% Jeffreys intervals on γobs propagated to k . Dashed lines are linear fits k⋆≈max(0,fg+a) ; the dotted line k=g marks full integration; the solid black line marks the random-guess boundary, below which Dk falls to chance and k⋆ is unidentified. 4B-Instruct without tools is omitted: its accuracy was indistinguishable from chance at every depth. Without tools, slopes are below 1 ( f=0.71 – 1.04 ) so the gap between k⋆ and g grows with depth, and accuracy correspondingly diverges from full-integration performance. 4B-Thinking’s slope of 1.04 is misleading: at small g , partial integration can yield correct answers by chance, so k⋆ tracks g until g≥15 when full integration becomes necessary and accuracy collapses, widening the band to the random-guess boundary. With tools, slopes are essentially 1 ( 0.96 – 1.08 ) and bands stay tight, so the error rate is independent of problem depth.
A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors. When sampling fails, a common fix shows the generator the gold answer and asks it to write a chain that reaches that answer. We show that this second step degrades the training data in a way that correctness filtering cannot catch. We run a controlled experiment that fixes the generator, the problem set, and the correctness filter, and varies only whether the chain is generated under answer-conditioning, the gold answer shown with a request to reach it. Training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy. The loss grows with difficulty, reaching as much as about 27 points on the hardest competition problems. The mechanism is legible in the chains themselves, which rationalize backward from the shown answer instead of deriving it, with the early final-answer statement as the measurable symptom. The harm is a property of the data rather than the generator, read off unlabeled generations before any fine-tuning, ordering the penalty across eight thinking models from four families, and transferring across teacher families. A prompt ablation localizes it to the rationalize-toward instruction rather than the answer's bare visibility. The practical takeaway is to generate answer-blind, because no correctness filter can see this damage in the data.
Jungseob Lee, Seungyoon Lee, Suhyune Son +4
1Korea University · 2Zoom Communications · 3Yonsei University
Large language models (LLMs) have become increasingly used for various tasks, often coupled with Chain-of-Thought (CoT) prompting to boost accuracy. Recent work has shown that high label-prediction accuracy does not guarantee correct intermediate reasoning, and the causes of reasoning flaws vary from sample to sample, yet existing remedies either focus on a single domain or assume that one flaw type applies uniformly across samples. A simple mitigation method is to provide the model with the correct answer, but we show that this yields no consistent improvement in reasoning quality. This indicates that the problem cannot be fixed by LLMs' awareness of answers, and must instead be addressed through the structure of reasoning. Motivated by this, we propose CRAFT (Consensus Reasoning-knowledge-graph Aggregation for Flaw-aware Trace synthesis), which aggregates the consensus components shared across multiple candidate reasoning traces to synthesize improved ones. CRAFT consistently improves label-prediction accuracy on both logical and mathematical reasoning benchmarks, outperforming most baselines, while its post-processed traces achieve higher quality under fine-grained benchmark evaluation.
Zipeng Ling, Shuliang Liu, Seonil Son +4
The Hong Kong University of Science and Technology (Guangzhou) · University of Alberta · Alberta Machine Intelligence Institute +3
When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct answers even when native sequence scoring completely collapses due to structural biases. To test whether instance-specific logic survives this collapse, we introduce a diagnostic protocol using a minimal, target-label-free additive correction. Fitting just two parameters on as few as 25 unlabeled examples recovers 9--34 accuracy points for Qwen3.5 models, transferring successfully to OLMo-2-1B and Llama-3.1-8B. Crucially, these recovered decisions persist on hard instances unresolved by simple lexical overlap and significantly exceed count-preserving permutation baselines. Our results show that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.
Qiyao Yan, Chenpeng Wang, Liangming Pan
State Key Laboratory of Multimedia Information Processing, Peking University · School of Computer Science, Peking University · YiXin-AILab, YIXIN, Beijing, China +1