Authors: Trevor McCourt, Ila R. Fiete, Isaac L. Chuang
Organizations: Department of Electrical Engineering and Computer Science, MIT · Department of Brain and Cognitive Sciences, McGovern Institute, MIT · Department of Physics, MIT
Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of training runs on simulated faulty digital hardware quantify this trend and suggest that models learn to compute within "good" error-correcting codes, whose relative overhead remains finite no matter how large the model gets. This finding leads us to conjecture that appropriately trained LLMs may be formally fault-tolerant; if true, running AI inference on low energy, faulty hardware may be a path to substantial energy savings over the status quo.
Figures & tables
Figure 1: The emergence of fault tolerance in AI models as they grow. Each curve shows how much useful capacity a language model retains when trained and run on hardware that introduces errors into arithmetic operations with probability p (darker, more faults). Such a fault-hardened model with N weights performs as well as a conventional model with η weights run on clean hardware, so the useful capacity η/N is the fraction of the model doing useful work rather than correcting errors (Eq. 1 ; unlimited training data). Capacity first falls with size, then recovers beyond a critical size N∗ , driven by a term in the scaling law that fault-free models lack (Eq. 4 ). This is reminiscent of “good” error-correcting codes, whose overhead stays a fixed fraction at any scale, and we conjecture by analogy that η/N approaches a nonzero constant whenever p lies below a threshold (Eq. 3 ). Solid curves and shaded bands show measured scaling-law fits and their bootstrap interquartile ranges; dashed curves show possible extrapolations, which depend on how much faults raise the best performance a model can ever attain (Eqs. 9 and 10 ).
Figure 2: Training under faults changes how language models scale. Fault-hardened transformers obey a scaling law (Eq. 4 ) with an extra coefficient compared to the standard form ( α2 ); each point fits all models trained at a fault rate p (Methods). a , Model-size coefficients. α1 , the typical model-size coefficient, shrinks as p increases. α2 , the curvature term absent from standard laws (pinned at zero for p=0 ), grows with p and drives the recovery of capacity beyond N∗ in Fig. 1 . b , Data coefficients. β1 , the rate at which loss falls with training data, decreases as p increases. c , Cost of robustness. Tokens D a fault-hardened model needs to match a fault-free model trained on δ tokens (Eq. 6 ); gray shading marks extrapolation beyond our largest runs. Error bars, 95% confidence intervals; bands in c , interquartile ranges; both from 1,500 bootstrap resamples per p cohort.
Figure 3: Inference-time faults are fatal to models trained without them. a, The divergence of fault-blind model scaling under inference-time errors. Conventionally, we expect language model performance to increase monotonically with D . However, when peval>0 , fault-blind models do not always exhibit this crucial behavior: performance scaling with D weakens as peval increases, and past a threshold, performance can actually decrease with D . Shaded regions indicate the standard error of the reported values, estimated from 8 independently trained models for each value of D . b, Error-induced heating of model logits. Part of the degradation caused to models by inference-time errors is heating of the output logits: fault-blind model predictions rise in temperature rapidly with peval , more so at large D , while fault-hardened models resist the heating. c, Error-induced scrambling of model logits. Even after removing the temperature contribution, fault-blind models lose capacity rapidly as peval rises above 0 , faster at larger D ; fault-hardened models maintain almost all capacity for peval<ptrain , independent of D . Bands in (b) and (c) indicate the IQR of the reported values, taken over the different trained models and contexts. d, Faulted predictions plotted against fault-free predictions, overlaid for many validation-set inputs. A temperature shift skews the logit blob: the ptrain=0 heating is obvious, along with substantial randomization, while the ptrain=peval logits are barely displaced.
Figure 4: Error-detecting hardware turns arbitrary arithmetic faults into dropped connections. a, The operations that faults targeted in our experiments. We injected errors into the matrix multiplications that drive the attention and feed-forward operations (orange) in a Llama-2-style transformer. b, The fraction of fault-targeted operations vs model size. Matrix multiplications (matmuls) account for almost all of the floating-point operations (FLOPs) required to run transformers; for large models, more than 99% of the FLOP budget. c, Matrix products reduce to dot products. In hardware, each matmul breaks down into a series of dot products u⋅v . d, The error model. A chain of fused multiply–add (FMA) cells accumulates each dot product. Operands carry a redundant residue (their remainder modulo small bases), and after every k=4 cells a cheap modulus check (%) tests the running sum for consistency. A failed check discards those k products and carries the previous partial sum forward, equivalent to zeroing a block of k elements in one operand; blocks fail independently with probability p . Detection thus converts hardware errors of arbitrary magnitude into well-behaved block drop-connect, leaving correction to the model.
Figure S1: Retained-capacity signals are visible in the raw data. (a) Subtracting the D -dependence from the raw data. The left panel shows the raw experimental data; the right panel shows the same data after LD has been subtracted from each loss value. Essentially all D -dependence is removed from the data, as far as is visible within run-to-run variation. We show the subtraction for a single value of p ; the effective datasets for all values of p are given in a notebook included in the code accompanying this manuscript [ 63 ] . (b) Retained-capacity signals in the LD -subtracted dataset. Once every datapoint has been brought to convergence by removing the D -dependence, the standard Chinchilla law fitted to the p=0 cohort can be used to compute η for each point in the data.
Figure S2: Residuals of the scaling law fit (a) The log-residuals of our scaling law fit ( L^ ) vs measured loss ( L ), for all measured values of p . The scale of the residuals is near what would be expected due to run-run variation alone (indicated by the gray band). (b) The loss predicted by our model vs the measured loss. The black line indicates a perfect fit.
Figure S3: Residuals in the scaling law fit without the curvature term Residuals in a fit of the standard Chinchilla form to our datasets. The curvature term is clearly needed to represent the trends in the raw data.
Figure S4: Pairwise correlations between scaling law parameters Pairwise scatter plots that identify correlations between scaling law parameters. A small subset of the complete bootstrap dataset is shown to avoid overcrowding the plots with points.