cs.LGOct 4, 2026

Expanding LLM Reasoning

Authors: Rian Atri, Evan Luo

Organizations: Keiji AI · University of California, Berkeley

Abstract

Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and measure it at every eligible step for nine models on six benchmarks (41 model and benchmark cells). Restart position matters: steps selected on one set of continuations beat uniform placement when scored on disjoint ones, in held-out audits on 5, 16, and 38 cells (+4.25 points [+2.51, +6.63] in a fresh five-cell audit). A fixed rule that restarts from the last eligible steps, always-last, is a strong baseline: our learned router beats uniform placement but shows no detected gain over it, and on DeepSeek-R1-Distill-Qwen-14B/MATH-500 always-last exceeds the exact self-consistency frontier at matched aggregate generated output by +0.052 [+0.008, +0.098], using 0.774x the aggregate generated output of four-sample self-consistency. Cross-fitted oracle selection still finds held-out headroom beyond declared positional classes, a target for future selectors. Finally, breaking step-label ties by earliest index flips the sign of a pointwise selector's gain over uniform placement in every seed of a five-seed diagnostic with four rollouts per step; randomized ties remove the bias.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm

    Sep 28, 2025Kaisen Yang, Tinghe Zhang, Rushi Shah +4LLM Inference EfficiencyLLM Planning

  2. Learning to Reason Efficiently with A* Post-Training

    May 23, 2026Andreas Opedal, Francesco Ignazio Re, Abulhair Saparov +3Process Reward ModelsRL for Language Model Reasoning