cs.DCSep 30, 2026

Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware

Authors: Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim

Organizations: The University of Texas at Austin Austin, Texas, USA

Abstract

Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training. We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two-tier erasure protection and chunk-level transactional updates to keep the retained model state recoverable and consistent, and reclaim the shadow state when active training needs its GPU memory. Evaluation on 6- and 72-GPU NVIDIA A100 clusters shows that Leto recovers 3.6--6.5×\times faster than the best-performing checkpointing baselines and improves productive training time by up to 13.7 percentage points. Large-scale simulation shows over 95% productive training time on a 131,072-GPU cluster.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint

    Jul 2, 2026Haotian Xie, Junlin Chen, Mingkai Zheng +2Large Language Model TrainingIntermediate Checkpoints

  2. ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload

    May 11, 2026Ziyue Liu, Zhengyang Wang, Ruijie Zhang +7Large Language Model TrainingFault Tolerance

  3. TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training

    May 18, 2026Shujie Han, Feng Jiang, Patrick P. C. Lee +5Model CheckpointsLarge Language Model Training