Authors: Trevor McCourt, Ila R. Fiete, Isaac L. Chuang
Organizations: Department of Electrical Engineering and Computer Science, MIT · Department of Brain and Cognitive Sciences, McGovern Institute, MIT · Department of Physics, MIT
Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of training runs on simulated faulty digital hardware quantify this trend and suggest that models learn to compute within "good" error-correcting codes, whose relative overhead remains finite no matter how large the model gets. This finding leads us to conjecture that appropriately trained LLMs may be formally fault-tolerant; if true, running AI inference on low energy, faulty hardware may be a path to substantial energy savings over the status quo.
Figures & tables
Figure 1: The emergence of fault tolerance in AI models as they grow. Each curve shows how much useful capacity a language model retains when trained and run on hardware that introduces errors into arithmetic operations with probability p (darker, more faults). Such a fault-hardened model with N weights performs as well as a conventional model with η weights run on clean hardware, so the useful capacity η/N is the fraction of the model doing useful work rather than correcting errors (Eq. 1 ; unlimited training data). Capacity first falls with size, then recovers beyond a critical size N∗ , driven by a term in the scaling law that fault-free models lack (Eq. 4 ). This is reminiscent of “good” error-correcting codes, whose overhead stays a fixed fraction at any scale, and we conjecture by analogy that η/N approaches a nonzero constant whenever p lies below a threshold (Eq. 3 ). Solid curves and shaded bands show measured scaling-law fits and their bootstrap interquartile ranges; dashed curves show possible extrapolations, which depend on how much faults raise the best performance a model can ever attain (Eqs. 9 and 10 ).
Figure 2: Training under faults changes how language models scale. Fault-hardened transformers obey a scaling law (Eq. 4 ) with an extra coefficient compared to the standard form ( α2 ); each point fits all models trained at a fault rate p (Methods). a , Model-size coefficients. α1 , the typical model-size coefficient, shrinks as p increases. α2 , the curvature term absent from standard laws (pinned at zero for p=0 ), grows with p and drives the recovery of capacity beyond N∗ in Fig. 1 . b , Data coefficients. β1 , the rate at which loss falls with training data, decreases as p increases. c , Cost of robustness. Tokens D a fault-hardened model needs to match a fault-free model trained on δ tokens (Eq. 6 ); gray shading marks extrapolation beyond our largest runs. Error bars, 95% confidence intervals; bands in c , interquartile ranges; both from 1,500 bootstrap resamples per p cohort.
Figure 3: Inference-time faults are fatal to models trained without them. a, The divergence of fault-blind model scaling under inference-time errors. Conventionally, we expect language model performance to increase monotonically with D . However, when peval>0 , fault-blind models do not always exhibit this crucial behavior: performance scaling with D weakens as peval increases, and past a threshold, performance can actually decrease with D . Shaded regions indicate the standard error of the reported values, estimated from 8 independently trained models for each value of D . b, Error-induced heating of model logits. Part of the degradation caused to models by inference-time errors is heating of the output logits: fault-blind model predictions rise in temperature rapidly with peval , more so at large D , while fault-hardened models resist the heating. c, Error-induced scrambling of model logits. Even after removing the temperature contribution, fault-blind models lose capacity rapidly as peval rises above 0 , faster at larger D ; fault-hardened models maintain almost all capacity for peval<ptrain , independent of D . Bands in (b) and (c) indicate the IQR of the reported values, taken over the different trained models and contexts. d, Faulted predictions plotted against fault-free predictions, overlaid for many validation-set inputs. A temperature shift skews the logit blob: the ptrain=0 heating is obvious, along with substantial randomization, while the ptrain=peval logits are barely displaced.
Figure 4: Error-detecting hardware turns arbitrary arithmetic faults into dropped connections. a, The operations that faults targeted in our experiments. We injected errors into the matrix multiplications that drive the attention and feed-forward operations (orange) in a Llama-2-style transformer. b, The fraction of fault-targeted operations vs model size. Matrix multiplications (matmuls) account for almost all of the floating-point operations (FLOPs) required to run transformers; for large models, more than 99% of the FLOP budget. c, Matrix products reduce to dot products. In hardware, each matmul breaks down into a series of dot products u⋅v . d, The error model. A chain of fused multiply–add (FMA) cells accumulates each dot product. Operands carry a redundant residue (their remainder modulo small bases), and after every k=4 cells a cheap modulus check (%) tests the running sum for consistency. A failed check discards those k products and carries the previous partial sum forward, equivalent to zeroing a block of k elements in one operand; blocks fail independently with probability p . Detection thus converts hardware errors of arbitrary magnitude into well-behaved block drop-connect, leaving correction to the model.
Figure S1: Retained-capacity signals are visible in the raw data. (a) Subtracting the D -dependence from the raw data. The left panel shows the raw experimental data; the right panel shows the same data after LD has been subtracted from each loss value. Essentially all D -dependence is removed from the data, as far as is visible within run-to-run variation. We show the subtraction for a single value of p ; the effective datasets for all values of p are given in a notebook included in the code accompanying this manuscript [ 63 ] . (b) Retained-capacity signals in the LD -subtracted dataset. Once every datapoint has been brought to convergence by removing the D -dependence, the standard Chinchilla law fitted to the p=0 cohort can be used to compute η for each point in the data.
Figure S2: Residuals of the scaling law fit (a) The log-residuals of our scaling law fit ( L^ ) vs measured loss ( L ), for all measured values of p . The scale of the residuals is near what would be expected due to run-run variation alone (indicated by the gray band). (b) The loss predicted by our model vs the measured loss. The black line indicates a perfect fit.
Figure S3: Residuals in the scaling law fit without the curvature term Residuals in a fit of the standard Chinchilla form to our datasets. The curvature term is clearly needed to represent the trends in the raw data.
Figure S4: Pairwise correlations between scaling law parameters Pairwise scatter plots that identify correlations between scaling law parameters. A small subset of the complete bootstrap dataset is shown to avoid overcrowding the plots with points.
Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training. We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two-tier erasure protection and chunk-level transactional updates to keep the retained model state recoverable and consistent, and reclaim the shadow state when active training needs its GPU memory. Evaluation on 6- and 72-GPU NVIDIA A100 clusters shows that Leto recovers 3.6--6.5× faster than the best-performing checkpointing baselines and improves productive training time by up to 13.7 percentage points. Large-scale simulation shows over 95% productive training time on a 131,072-GPU cluster.
Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim
The University of Texas at Austin Austin, Texas, USA
State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack. Existing fault-tolerance mechanisms either impose non-trivial overhead during failure-free execution or suffer from prolonged recovery latency, particularly under scenarios where a small subset of compute nodes experience permanent failures. %The tradeoff between failure-free overhead and recovery latency forms a space forms a Pareto frontier We present PHOENIX to simultaneously address both optimization objectives. PHOENIX incorporates a fault-tolerance mechanism that restores LLM training via hot-swapping, namely by replacing failed nodes with spare nodes without terminating the complete job. The hot-swapping of PHOENIX is enabled by two ideas: First, it exploits an off-critical-path in-memory checkpointing mechanism for spatial redundancy. Second, it introduces a communicator reconstruction protocol that replaces failed nodes with spare nodes at runtime. PHOENIX efficiently overlaps the in-memory checkpointing with computation, thus introducing zero overhead during error-free execution. Upon permanent node failures, PHOENIX can rebuild memory states with minimal recomputation by leveraging in-memory checkpoints. We evaluate PHOENIX across scales (up to 512 NVIDIA A100 GPUs) and LLMs (up to 65B parameters), and observe zero checkpoint overhead with hot-swapping recovery completing in under 40 seconds. These results show that PHOENIX simultaneously achieves both zero-overhead error-free execution and extremely low recovery cost.
Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve state-of-the-art performance on unperturbed inputs using standard predictive metrics, their performance on perturbed inputs, measured using reliability metrics, remains underexplored, despite its importance for reliable deployment. To address this gap, we first conduct a comprehensive reliability evaluation of quantized LLMs consisting of three key components: (1) Uncertainty: We assess the trustworthiness of LLMs quantized to 2, 3, 4, and 8 bits using six different quantization methods, employing established uncertainty metrics. (2) Calibration: We assess how well-calibrated the uncertainty estimates of quantized models are across model scales and bit precisions. (3) Robustness: We design character-level and word-level input perturbations to evaluate the reliability of quantized models under semantically-preserving variations in the inputs that arise in real-world applications. Second, we characterize how reliability scales with the total number of model bits. Our study reveals that while the performance scales monotonically with the total number of bits, the reliability scalings are nonlinear. A reliability peak occurs for 4-bit quantized models, indicating that quantizing moderately sized models offers the best reliability-efficiency trade-off. Additionally, our empirical findings reveal that quantization enhances the robustness of LLMs to natural input perturbations.
Sirine Ayadi, Sándor Daróczi, Stephan Günnemann +1
School of Computation, Information and Technology, Technical University of Munich Munich Data Science Institute · Pruna AI