A Tilted Bowl Is Not a Slippery Slope: Compressing Looped Models
Organizations: Carnegie Mellon University · ML Collective
Abstract
Looped models reason by applying the same block of weights many times, so compressing that block saves memory traffic on every loop. Compressed looped models, however, often collapse, and the collapse is usually blamed on rounding error that accumulates from loop to loop. In this work we test that account on more than 30 models from five families and find, to our surprise, that it holds only for loops that never settle. When a loop settles, a fixed rounding error does not accumulate. It moves the point where the loop settles, much as tilting a bowl moves where a ball comes to rest, and the answer is lost only when the shift is larger than the readout tolerates. This picture lets us predict which models fail from a single label-free measurement, and it tells us why failed models recover: their loops still settle, so a few final loops with 8-bit weights bring the answer back. Motivated by these findings, we build a controller that stops when the model's halting head fires and then finishes with 8-bit loops. On Sudoku-Extreme and Maze-Hard it beats fixed-depth inference by up to 15 points under a third of the weight traffic.
Figures & tables
| Family (models) | Models | Task | What loops |
| Puzzle reasoners (24) | TRM with an MLP or an attention mixer, HRM ( Jolicoeur-Martineau, 2025 ; Wang et al., 2025 ) ; 19 TRMs we train at several widths and seeds | Sudoku-Extreme, Maze-Hard | one block, cycles per supervision step |
| Recurrent convnets (1) | DeepThinking ( Bansal et al., 2022 ) | mazes of size 13 and 33 | a recall block, input re-injected |
| Implicit models (3) | MDEQ-Small, MDEQ-XL ( Bai et al., 2020 ) ; DEQ-Transformer ( Bai et al., 2019 ) | ImageNet, WikiText-103 | a solver, run to a fixed point |
| Looped language models (4) | Huginn-3.5B ( Geiping et al., 2025 ) ; Ouro-1.4B, Ouro-2.6B ( Zhu et al., 2025 ) ; Recurrent-OLMo-2 ( McLeish et al., 2025 ) | GSM8K | a core block, 4–32 recurrences; full precision is bf16 |
| Diffusion, the boundary (1) | DDPM UNet, sampled by DDPM ( Ho et al., 2020 ) and DDIM ( Song et al., 2021 ) | CIFAR-10 | the denoiser; DDPM draws fresh noise each step, DDIM is deterministic |
| w4c | w3a | |||||
| Model (loops) | Settles | bf16 | alone | +precise | alone | +precise |
| Huginn-3.5B (32) | ✓ | 42 | 33 | 37 | 3 | 40 |
| Recurrent-OLMo-2 (32) | ✓ | 41 | 35 | 46 | 2 | 45 |
| Ouro-1.4B (4) | ✗ | 73 | 7 | 66 | 0 | 0 |
| Ouro-2.6B (4) | ✗ | 81 | 78 | 74 | 0 | 50 |
| Method | Accuracy (%) | Cost | Finishing steps |
| Fixed depth | 61.1 | 55.0 | 0 |
| + stop at the halting head | 63.0 | 23.4 | 0 |
| + 8-bit finishing = controller | 75.8 | 16.1 | 1.13 |
| controller fixed depth | [ , ] | ||
| + latent sampling (add-on) | 78.7 | 28.4 | 1.45 |
| Controller with full-precision finishing | 75.7 | 25.8 | 1.20 |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | # | What loops | Default depth | Task, metric | |
| TRM, MLP mixer (released) | 1 | one 2-layer block, reused for the and updates; token-mixing SwiGLU | , , 16 supervision steps | Sudoku-Extreme (97 positions), exact | 1000; 1800 (depth to 512) |
| TRM, attention mixer (released) | 2 | as above, self-attention with RoPE; one checkpoint per task | as above | Sudoku-Extreme; Maze-Hard (916 positions), exact and valid path | 1000 / 500 |
| HRM (released) | 2 | separate high- and low-level modules (4 layers each); one checkpoint per task | , , 16 supervision steps | Sudoku-Extreme, Maze-Hard | 1000 / 500 |
| Our TRMs, MLP mixer | 10 | as TRM; width 256 (seeds 0–2), 512 (seeds 0–5), 768 (seed 0) | as TRM | Sudoku-Extreme, exact | 1000 |
| Our TRMs, attention mixer | 9 | as TRM; width 256 (seeds 0–2), 512 (seeds 0–5) | as TRM | Sudoku-Extreme, exact | 1000 |
| DeepThinking | 1 | recall block, input concatenated at every iteration (width 128) | 300 iterations; swept to 1000 | mazes of size 13 and 33, exact | per size |
| Dense | 25% pruned | 50% pruned | Distilled | |||||
| Task | exact | cell | exact | cell | exact | cell | exact | cell |
| ARC | 36.0 | 87.6 | 0.0 | 0.3 | 0.0 | 0.6 | – | – |
| Maze | 86.8 | 99.5 | 0.0 | 86.4 | 0.0 | 0.0 | 0.0 | 87.3 |
| Sudoku | 69.1 | 87.5 | 0.0 | 50.0 | 0.0 | 37.8 | 0.0 | 54.8 |
| Compressed format | w4 | w3 | w3g32 | w2 | w2g32 | w8 |
| 4 | 3 | 3.5 | 2 | 2.5 | 8 | |
| (full-precision loop) | 8 | 10.7 | 9.1 | 16 | 12.8 | 4 |
| Endpoints | Trend | |||||
| Run (series) | ||||||
| A100 (22) | 1 / 16 / 5 | 0 / 18 / 4 | 3 / 14 / 5 | 0 / 15 / 7 | 0 / 19 / 3 | 5 / 12 / 5 |
| L40S replicate (23) | 0 / 15 / 8 | 0 / 20 / 3 | – | 0 / 18 / 5 | 0 / 20 / 3 | – |
| L40S independent (23) | 2 / 18 / 3 | 0 / 21 / 2 | – | 1 / 18 / 4 | 0 / 21 / 2 | – |
| Model | Format | |||
| TRM Sudoku, MLP | w4c | +1.1 [ 1.9, +4.1] = | 1.2 [ 3.3, +1.1] = | +1.4 [ 1.9, +4.8] = |
| w4g32 | 0.2 [ 2.4, +2.1] = | 1.2 [ 3.3, +0.9] = | 1.3 [ 4.4, +1.8] = | |
| W4A4 g32 | +5.0 [+2.2, +7.9] | 0.3 [ 2.5, +1.9] = | 7.5 [ 10.7, 4.5] | |
| W4A4 MXInt4 | +7.5 [+4.1, +10.9] | +0.1 [ 2.3, +2.3] = | 11.3 [ 14.4, 7.9] | |
| W8A8 per-tensor | +13.8 [+10.8, +16.9] | +2.5 [+0.2, +5.0] | +7.3 [+3.8, +10.8] | |
| TRM Sudoku, attn. | w4c | +0.7 [ 1.0, +2.4] = | +0.3 [ 2.0, +2.6] = | +4.7 [+2.3, +7.1] |
| Noise / format | Cosine to full precision | Converged / median steps | ||||
| fp32 | 72.2 | 73.9 | 73.8 | 73.8 | 1 | 99.7% / 24 |
| 0.05 | 71.4 | 72.9 | 73.1 [72.7, 73.7] | 73.0 | 0.978 | 99.7% / 24 |
| 0.1 | 67.3 | 69.0 | 68.9 [67.1, 70.7] | 68.9 | 0.923 | 99.8% / 24 |
| 0.2 | 46.0 | 49.1 | 49.1 [47.1, 52.9] | 49.1 | 0.779 | 99.8% / 25 |
| 0.3 | 12.6 | 14.6 | 14.8 [10.5, 18.4] | 14.8 | 0.614 | 99.8% / 25 |
| 0.5 | 0.3 | 0.2 | 0.3 | 0.3 | 0.334 | 99.5% / 26 |
| Noise / format | Residual, | Cosine, | |||
| fp32 | 91.2 | 24.48 | 23.14 | 2.5e-2 | 1 |
| 0.1 | 92.3 | 25.23 | 23.84 [23.82, 23.86] | 2.4e-2 | 0.981 |
| 0.3 | 98.5 | 34.59 | 32.65 [32.63, 32.70] | 2.3e-2 | 0.846 |
| 0.5 | 219 | 124.7 | 127.1 [124.2, 132.0] | 2.9e-2 | 0.55 |
| 0.7 | 923 | 1071 | 1182 [1142, 1204] | 6.1e-2 | 0.29 |
| w8c (0.008) | 90.2 | 24.48 | 23.14 | 2.5e-2 | 1.000 |
| DDIM steps | DDPM steps | |||||||
| / format | 10 | 25 | 50 | 100 | 50 | 100 | 250 | 1000 |
| 0.01 | 7.4 | 6.6 | 6.1 | 6.7 | 1.5 | 2.1 | 2.6 | – |
| 0.03 | 26.5 | 17.0 | 17.2 | 49.8 | 6.2 | 7.1 | 7.9 | – |
| 0.05 | 45.7 | 21.2 | 61.4 | 189.5 | 12.7 | 13.3 | 13.8 | – |
| 0.1 | 95.6 | 136.5 | 345.8 | 397.6 | 43.2 | 42.8 | 42.1 | 46.1 |
| 0.3 | 386 | 358 | 356 | 312 | 213 | 255 | 290 | – |
| Clipping | Quantized layers | Maze exact | Maze valid path | Sudoku-MLP exact | Maze / MLP |
| max over calls (ours) | all linears | 0.0 [0.0–0.0] | 0.0 | 0.0 [0.0–0.0] | 0.293 / 0.466 |
| max over calls | block only | 0.0 [0.0–0.0] | 0.0 | 0.0 [0.0–0.0] | 0.293 / 0.465 |
| 99.99th percentile | all linears | 82.2 [82.2–82.2] | 98.4 | 0.8 [0.4–1.0] | 0.145 / 0.270 |
| 99.99th percentile | block only | 82.3 [82.2–82.4] | 98.4 | 0.8 [0.5–1.0] | 0.144 / 0.269 |
| Model | Error | ( ) | Gap 16 / 64 / 512 | Per-draw regimes | |
| TRM MLP | fresh | 0.15 (0.89) | / / | [ , ] | full catch-up 5 |
| fresh | 0.2 (0.87) | / / | [ , ] | full 1, partial 4 | |
| fresh | 0.25 (0.70) | / / | [ , ] | constant 5 | |
| fixed | 0.08 (0.88) | / / | [ , ] | constant 5 | |
| fixed | 0.1 (0.85) | / / | [ , ] | constant 5 | |
| fixed | 0.12 (0.72) | / / | [ , ] | widening 5 |
| Variant | 0.1 | 0.15 | 0.2 | 0.3 | 0.5 | ||
| an fresh, zero-mean | 70.1 [70–70] | 67.4 [67–68] | 62.9 [63–63] | 32.9 [32–34] | 2.3 [2–2] | 0.287–0.293 | 0.080–0.082 |
| af fixed, zero-mean | 66.5 [66–67] | 58.6 [57–60] | 43.8 [41–46] | 10.4 [8–13] | 0.3 | 0.216–0.231 | 0.061–0.064 |
| ab fixed, sign-consistent | 64.6 [62–67] | 46.5 [35–56] | 22.9 [10–31] | 3.7 [0–6] | 0.0 | 0.149–0.190 | 0.094–0.124 |
| wn weight, zero-mean | 64.9 [63–66] | 26.5 [24–31] | 5.5 [4–7] | 1.0 | 0.0 | 0.135–0.144 | 0.038–0.042 |
| Variant | 0.5 | 0.7 | 1.0 | 1.5 | exact | valid path | |
| an fresh, zero-mean | 87.5 [87–88] | 85.5 [85–87] | 23.3 [23–24] | 0.0 | 0.897–0.902 | 1.20–1.21 | 0.198–0.199 |
| af fixed, zero-mean | 77.7 [70–82] | 34.9 [20–48] | 0.0 | 0.0 | 0.60–0.73 | 0.78–0.83 | 0.13–0.16 |
| ab fixed, sign-consistent | 86.1 [84–88] | 79.5 [72–84] | 9.3 [0–25] | 0.0 | 0.82–0.90 | 0.83–0.91 | 0.18–0.20 |
| wn weight, zero-mean | 87.7 [87–88] | – | 86.1 [85–87] | 27.1 [0–81] | 1.24–1.73 | 1.73–1.75 | 0.30–0.44 |
| Model / format (regime) | Mixer-path error, steps 1 / 4 / 16 / 64 / 512 | Whole-step error | Entropy | Top probability |
| attention, full precision | 0 | 0 | 0.694 0.643 | 0.21 0.26 |
| attention w3c (catch-up) | 0.210 / 0.281 / 0.287 / 0.290 / 0.291 | 0.636 0.289 | 0.719 0.646 | 0.19 0.26 |
| attention w3a (catch-up) | 0.170 / 0.232 / 0.237 / 0.239 / 0.240 | 0.444 0.229 | 0.703 0.643 | 0.20 0.26 |
| attention w3g32 (catch-up) | 0.136 / 0.181 / 0.187 / 0.189 / 0.190 | 0.310 0.168 | 0.696 0.638 | 0.21 0.26 |
| attention w4t (constant) | 0.169 0.240 | 0.503 0.229 | 0.711 0.651 | 0.20 0.25 |
| MLP w4c (constant) | 0.086, flat | 0.520 0.220 | – | – |
| Clause | Bar | Outcome | Verdict |
| (A1) collapsed models with | 80% | 12% | fails |
| (A2) surviving models with | 80% | 91% | holds |
| (B1) collapsed models with | 80% | 88% | holds |
| (B2) zero-parameter width, new seeds | bias in [0.9, 1.1] | 8.9 (bias 8.25 / 9.60) | fails |
| one fitted scale | (reported) | 1.57 (bias 1.45 / 1.68) | – |
| reference: , | – | 1.39, 1.16 | – |
| ID | Order | Prediction and bar | Outcome |
| P1 | 1 | A compressed model collapses iff its settled state does not return under the full-precision loop: rule “collapse iff median ” balanced accuracy ; AUROC( ) AUROC( ); leave-one-family-out ; Spearman(retained, return) . | Falsified on both platforms. L40S: balanced accuracy 0.63; AUROC 0.70 vs 0.98; LOFO 0.72 vs 0.94; 29/64 collapsed models return from the grid cap. A100: 0.61; AUROC 0.60 vs 0.97; LOFO 0.78 vs 0.91; 17/28 at the cap. |
| P2 | 1 | Finishing law: gain of full-precision finishing = return rate gap; supported if pred obs on 80% of matched models. | Confirmed : 20/20 matched models, mean pred obs 1.52 points, (L40S). |
| P3 | 1 | Sampling: the halting head picks draws with larger (sign test in each model), and the best-of-16 gain peaks at intermediate return rates. | Falsified . The head picks larger- draws in HRM w3c (57/18) and MLP w4c (15/1), not attention w3c (64/67); the gain is largest at the lowest return rate. The head does pick correct draws (picked-correct vs median draw: attention 68/0, HRM 89/0, MLP 62/0 on the L40S platform; 72/0, 83/1, 68/1 on the A100 platform). |
| P4 | 1 | Radius law: median at noise 0.1 removes the tolerance law’s mixer bias. | Falsified : , worse than the laws with and . |
| K1 | 2 | full-precision finishing steps recover 80% of the examples full precision solves, on collapsed models. | Falsified on the L40S platform, partly on the A100 platform; 8–16 steps needed (Table 19 ). |
| K2 | 2 | w8c finishing within 2 points of full-precision finishing on 90% of compressed models (at and 16). | Partly on the L40S platform (0.84 of models), confirmed on the A100 platform (0.93); median w8c fp32 0.003 (L40S), 0.000 (A100). |
| Platform | (full-precision finishing steps) | 1 | 2 | 4 | 8 | 16 |
| L40S | recovery (median model) | 0.28 | 0.50 | 0.71 | 0.83 | 0.89 |
| retained accuracy | 0.29 | 0.51 | 0.75 | 0.88 | 0.96 | |
| A100 | recovery (median model) | 0.32 | 0.67 | 0.80 | 0.87 | 0.88 |
| retained accuracy | 0.33 | 0.69 | 0.81 | 0.88 | 0.96 |
| Model | Family | Full precision | |||
| TRM Sudoku MLP | TRM | 0.715 | 0.14 | 1.74 | 0.24 |
| TRM Sudoku attention | TRM | 0.727 | 0.80 | 0.69 | 0.55 |
| TRM Maze attention | TRM | 0.875 | 1.71 | 0.32 | 0.55 |
| HRM Sudoku | HRM | 0.477 | 1.05 | 0.35 | 0.37 |
| HRM Maze | HRM | 0.748 | 1.71 | 0.23 | 0.40 |
| DeepThinking maze 13 | DeepThinking | 1.000 | 0.37 | 0.92 | 0.34 |
| Variant | Full precision | |||
| attention w512 s1 | 0.629 | 0.65 | 0.80 | 0.52 |
| attention w512 s2 | 0.534 | 0.60 | 0.85 | 0.51 |
| MLP w256 s0 | 0.599 | 0.14 | 1.46 | 0.21 |
| MLP w256 s1 | 0.623 | 0.13 | 2.05 | 0.27 |
| MLP w256 s2 | 0.640 | 0.13 | 1.30 | 0.17 |
| MLP w512 s0 | 0.836 | 0.16 | 1.50 | 0.25 |
| Seed | Full precision | w4t ( ) | predicted | measured | error |
| attention s3 | 68.6 | 27.9 (0.741) | 0.537 | 0.709 | 1.32 |
| attention s4 | 71.1 | 60.1 (0.840) | 0.563 | 0.578 | 1.03 |
| attention s5 | 64.3 | 57.7 (0.901) | 0.603 | 0.723 | 1.20 |
| MLP s3 | 82.1 | 3.9 (0.559) | 0.268 | 0.176 | 1.52 |
| MLP s4 | 81.3 | 6.1 (0.534) | 0.284 | 0.158 | 1.80 |
| MLP s5 | 82.4 | 7.8 (0.498) | 0.236 | 0.147 | 1.60 |
| Law | LOO median / worst | new seeds median / worst | bias MLP / attn / HRM | |
| 0.84 | 1.47 / 2.07 | 1.39 / 1.61 | 1.37 / 0.64 / 0.89 | |
| 0.91 | 1.35 / 1.96 | 1.16 / 1.29 | 1.19 / 0.76 / 0.98 |
| Model | implied room | pred/meas ( ) | ||
| LSQ-QAT s0 / s1 / s2 | 0.92 / 1.21 / 0.95 | 0.55 / 0.95 / 0.69 | 0.21 / 0.28 / 0.22 | 1.58 / 1.17 / 1.55 |
| fine-tune s0 | 1.57 | 1.09 | 0.21 | 1.57 |
| MLP w256 s0 / s1 / s2 | 1.29–2.05 | 1.05–1.86 | 0.19 / 0.27 / 0.17 | 1.73 / 1.22 / 1.97 |
| Model | Format | bf16 | compressed | +precise | return (teacher-forced) | return (answer) | ||
| Huginn | w8c | 42 | 36 | 41 | 1.000 | 0.25 | 100% | 80% |
| Huginn | w4g32 | 42 | 35 | 39 | 0.978 | 0.25 | 97% | 67% |
| Huginn | w4c | 42 | 33 | 37 | 0.929 | 0.25 | 91% | 55% |
| Huginn | w3a † | 42 | 3 | 40 | 0.782 | 0.25 | 85% | 47% |
| Huginn | w4t † | 42 | 2 | 38 | 0.696 | 0.25 | 89% | 54% |
| Recurrent-OLMo-2 | w8c | 41 | 43 | 43 | 1.000 | 0.25 | 100% | 86% |
| Loops | bf16 | w8c | w4g32 | w4c | w4t |
| 4 | 44.0 | 43.0 (1.00) | 41.0 (0.98) | 31.0 (0.942) | 0.0 (0.12) |
| 8 | 44.0 | 51.0 (1.00) | 40.0 (0.98) | 37.0 (0.947) | 0.0 (0.11) |
| 16 | 45.0 | 47.0 (1.00) | 41.0 (0.98) | 36.0 (0.947) | 0.0 (0.12) |
| 32 | 42.0 | 44.0 (1.00) | 42.0 (0.98) | 38.0 (0.946) | 0.0 (0.11) |
| Stopping rule | accuracy | Mean steps |
| Halting head, rule frozen per model | 8.15 | |
| Step-size threshold, frozen on full precision | 6.04 | |
| Step-size threshold, tuned per model | – | |
| Stop when the step size stops shrinking | 7.2 |
| Compressed format | w4 | w3 | w3g32 | w2g32 | w2 |
| 4 | 3 | 3.5 | 2.5 | 2 | |
| Full-precision finishing, | 8 | 10.7 | 9.1 | 12.8 | 16 |
| 8-bit finishing, | 2 | 2.67 | 2.29 | 3.2 | 4 |
| Model | Format | Fixed | Stop | Finish all | Controller | Low confidence | Stop, sample | +Sampling | |
| TRM-MLP | w3g32 | 9.1 | 59.1/32 | 61.5/22 | 74.2/51 | 74.2/51 | 74.1/32 | 71.2/52 | 76.4/56 |
| TRM-MLP | w4c | 8.0 | 77.5/64 | 77.5/19 | 76.5/48 | 77.5/19 | 76.5/18 | 82.6/36 | 82.6/36 |
| TRM-MLP | w4t | 8.0 | 7.0/64 | 7.0/64 | 67.3/42 | 67.3/42 | 66.8/40 | 7.7/64 | 67.3/42 |
| TRM-attn | w3c | 10.7 | 82.9/64 | 80.5/17 | 80.6/58 | 80.5/17 | 79.7/40 | 86.3/32 | 86.9/34 |
| TRM-attn | w4c | 8.0 | 83.8/64 | 83.7/15 | 79.4/47 | 83.7/15 | 79.8/15 | 86.8/32 | 86.8/32 |
| HRM | w2g32 | 12.8 | 20.8/64 | 20.5/62 | 37.0/52 | 37.0/52 | 34.6/55 | 26.0/64 | 37.0/52 |
| Full-precision finishing | 8-bit finishing | ||||||
| Model | Format | Head | Random | Finished | Head | Random | Finished |
| TRM-MLP | w3g32 | 74.1 | 69.2 [68.2, 70.1] | 0.64 | 74.8 | 69.4 [68.3, 70.4] | 0.64 |
| TRM-MLP | w4c | 76.5 | 74.1 [73.6, 74.6] | 0.27 | 76.8 | 74.2 [73.6, 74.7] | 0.27 |
| TRM-MLP | w4t | 66.8 | 66.8 [66.8, 66.8] | 1.00 | 67.6 | 67.6 [67.6, 67.6] | 1.00 |
| TRM-attn | w3c | 76.9 | 72.4 [71.6, 73.3] | 0.27 | 76.9 | 72.5 [71.6, 73.3] | 0.27 |
| TRM-attn | w4c | 79.4 | 78.7 [78.4, 79.0] | 0.21 | 79.6 | 78.7 [78.5, 79.1] | 0.21 |
| TRM-MLP w4c | TRM-attn w3c | HRM w3c | |
| Stochastic rounding, best of 16 (held out) | 93.7 | 94.2 | 79.2 |
| Full precision + latent noise, best-of-16 | 88.0 | 87.1 | 73.3 |
| Difference [95% CI] | [ , ] | [ , ] | [ , ] |
| Stochastic-rounding draws (rows 0–999) | 86.9 | 86.4 | 71.9 |
| Gaussian noise at the draws’ variance (rows 0–999) | 87.0 | 85.7 | 71.7 |