cs.AISep 22, 2026

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Authors: Minghui LiuThomas MagelinskiDehao YuanQi YuFurong Huang

Abstract

Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.

Explore similar work

CardsList
  1. Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

    Aug 13, 2025Xiaojun Wu, Xiaoguang Jiang, Huiyang Li +11Scaling Laws