cs.AIFeb 24, 2026

Tool Use Reduces Depth-Induced Collapse in OOD Reasoning

Authors: David Koplow, Tomer Galanti, Tomaso Poggio

Organizations: Massachusetts Institute of Technology · Texas A&M University

Abstract

Humans can apply ideas learned in one context to substantially different situations. We call this process of searching for and constructing novel recombinations of learned relationships to solve new problems \textit{out-of-distribution (OOD) reasoning}. The capacity for large language models (LLMs) to support OOD reasoning underpins proposals for generally intelligent systems. However, this property is challenging to measure because most problems admit many decompositions, some involving shallow subproblems and others involving subproblems that may have been memorized from the training data. This makes it difficult to determine how much compositional reasoning a model must actually perform. Uncertainty about training distributions, how to measure a datapoint's distance from a training distribution, and the exponential number of ways to decompose most tasks make it intractable to robustly measure any modern LLM's OOD-reasoning capacity on standard natural-language tasks. In this work, we introduce a benchmark in which a model progressively solves a Boolean circuit over GF(2)GF(2) from data. This benchmark is both minimal, isolating OOD reasoning from these confounding factors, and general, as any computable function can be represented as a sufficiently large GF(2)GF(2) polynomial. We find that standalone models' next-step accuracy collapses as depth grows. In contrast, tool use through the synthesis and execution of code prevents this collapse in both small and frontier LLMs. These results indicate that synthesizing tools play a crucial role in supporting OOD reasoning.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

    Jul 16, 2026Jungseob Lee, Seungyoon Lee, Suhyune Son +4LLM Reasoning StrategiesReasoning Skills

  2. Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

    Apr 15, 2026Zipeng Ling, Shuliang Liu, Seonil Son +4Reasoning TracesChain-of-Thought Reasoning

  3. Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

    Aug 31, 2026Qiyao Yan, Chenpeng Wang, Liangming PanLLM Reasoning StrategiesReasoning Benchmark