Enabling Dynamic Computation in Looped LMs
Organizations: Qualcomm AI Research · Johns Hopkins University
Abstract
Looped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens). However, state-of-the-art open Looped LMs trained with this dynamic computation capability (Ouro models) do not realize it in practice as each loop iteration (depth) requires its own level of KV-cache, necessitating all loop computations. Moreover, Ouro's early-exit prior is enforced on each token equally, which results in static lower-depth like processing of all tokens regardless of difficulty. In this work, we propose a simple "best-available" KV caching strategy that works out-of-the-box, creating a new frontier in the performance vs depth space. Our approach enables up to 30% reduction in FLOPs and KV memory while retaining full-depth performance, showing the true flexibility of Looped LMs. Furthermore, training looped LMs with awareness about this KV caching strategy improves performance and efficiency. Finally, we apply a small but effective fix to the early-exit prior enforcement objective that makes tokens exit at truly heterogeneous depths based on effort. Our findings are validated on Ouro models as well as smaller looped LMs pre-trained from scratch.
Figures & tables
| Dataset | Metric | Fixed-Depth | Dense-Ragged | ETE-Ragged | ||||
|---|---|---|---|---|---|---|---|---|
| depth / thr. | 2 | 3 | 4 | .3 | .5 | .3 | .5 | |
| Avg. depth | 2.00 | 3.00 | 4.00 | 2.39 | 3.05 | 2.63 | 3.21 | |
| Log-likelihood evaluation | ||||||||
| MMLU | accuracy | 60.4 0.4 | 66.7 0.4 | 67.5 0.4 | – | – | 61.9 0.4 | 66.2 0.4 |
| ARC-C | accuracy | 54.9 1.5 | 59.9 1.4 | 59.8 1.4 | – | – | 59.0 1.4 | 60.0 1.4 |
| HellaSwag | accuracy | 70.8 0.5 | 73.1 0.4 | 73.4 0.4 | – | – | 72.8 0.4 | 73.2 0.4 |
| Dataset | Metric | Fixed-Depth | Dense-Ragged | ETE-Ragged | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| depth / thr. | 2 | 3 | 4 | .3 | .45 | .5 | .3 | .45 | .5 | |
| Avg. depth | 2.00 | 3.00 | 4.00 | 2.24 | 2.45 | 2.83 | 2.30 | 2.60 | 3.00 | |
| AIME24 | pass@1 | 40.4 6.7 | 62.1 7.0 | 65.8 6.8 | 56.7 7.3 | 57.9 7.3 | 61.7 7.2 | 55.0 7.1 | 59.6 7.2 | 61.7 6.9 |
| pass@4 | 64.4 7.7 | 79.9 6.9 | 82.0 6.9 | 76.0 7.3 | 75.2 7.4 | 78.7 6.8 | 77.3 6.4 | 76.4 7.3 | 80.4 6.9 | |
| AIME25 | pass@1 | 31.7 6.7 | 49.2 7.7 | 47.9 7.9 | 45.4 7.2 | 45.4 7.4 | 52.9 7.2 | 46.2 7.2 | 49.2 7.3 | 52.1 7.5 |
| pass@4 | 49.0 8.5 | 65.6 8.2 | 62.1 8.5 | 67.7 7.6 | 64.8 7.9 | 70.8 8.1 | 66.9 8.0 | 68.7 7.9 | 69.2 8.0 | |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric | Measured | Comparator |
|---|---|---|
| Avg. decode depth | 2.252 | 4.000 (full depth) |
| Active decode row-depths | 18,373 | 32,640 (full depth) |
| Physical K/V rows | 25,621 | 40,016 (dense) |
| Nested depth passes | 823 | 574.2 (ideal ragged) |
| Benchmark | Examples | Shots | Metric |
|---|---|---|---|
| MMLU ( Hendrycks et al., 2021 ) | 14,042 | 5 | Accuracy |
| ARC-Challenge ( Clark et al., 2018 ) | 1,172 | 25 | Length-normalized accuracy |
| HellaSwag ( Zellers et al., 2019 ) | 10,042 | 10 | Length-normalized accuracy |
| WinoGrande ( Sakaguchi et al., 2021 ) | 1,267 | 5 | Accuracy |
| Benchmark | Examples | Shots | Max. tokens | Prompt |
|---|---|---|---|---|
| MMLU-Pro ( Wang et al., 2024 ) | 12,032 | 5 | 2,048 | Answer generation |
| BBH ( Suzgun et al., 2022 ) | 6,511 | 3 | 1,024 | Chain of thought |
| GSM8K ( Cobbe et al., 2021 ) | 1,319 | 3 | 1,024 | Chain of thought |
| MATH500 ( Lightman et al., 2024 ) | 500 | 5 | 2,048 | Answer generation |
| Benchmark | Examples | Rollouts | Max. tokens | Scoring |
|---|---|---|---|---|
| AIME 2024 | 30 | 8 | 32,768 | Qwen |
| AIME 2025 | 30 | 8 | 32,768 | Qwen |
| AIME 2026 | 30 | 8 | 32,768 | Qwen |
| GPQA-Diamond ( Rein et al., 2023 ) | 198 | 8 | 32,768 | Qwen |
| OlympiadBench ( He et al., 2024 ) | 581 | 4 | 32,768 | Qwen |
| HumanEval+ ( Chen et al., 2021 ) | 164 | 8 | 8,192 | EvalPlus ( Liu et al., 2023 ) |
| Benchmarks | Link | License |
|---|---|---|
| MMLU | HuggingFace | MIT |
| ARC-Challenge | HuggingFace | CC BY-SA 4.0 |
| HellaSwag | HuggingFace | MIT |
| Winogrande | HuggingFace | Apache 2.0 |
| MMLU-Pro | HuggingFace | MIT |
| BBH | HuggingFace | Apache 2.0 |
| Dataset | Link | License |
|---|---|---|
| FineWeb-Edu | HuggingFace | ODC-BY |
| AceReason-1.1-SFT | HuggingFace | CC BY 4.0 |
| OpenThoughts3 | HuggingFace | Apache 2.0 |
| Model | Link | License |
|---|---|---|
| Ouro-1.4B | HuggingFace | Apache 2.0 |
| Ouro-1.4B-Thinking | HuggingFace | Apache 2.0 |
| Ouro-2.6B | HuggingFace | Apache 2.0 |
| Ouro-2.6B-Thinking | HuggingFace | Apache 2.0 |
| Nanbeige-4.2-3B | HuggingFace | Apache 2.0 |
| Qwen-3.8-27B | HuggingFace | Apache 2.0 |
| Dataset | Metric | Fixed-Depth | Dense-Ragged | ETE-Ragged | ||||
|---|---|---|---|---|---|---|---|---|
| depth / thr. | 2 | 3 | 4 | .3 | .5 | .3 | .5 | |
| Avg. depth | 2.00 | 3.00 | 4.00 | 2.39 | 3.05 | 2.63 | 3.21 | |
| Log-likelihood evaluation | ||||||||
| MMLU | accuracy | 60.4 0.4 | 66.7 0.4 | 67.5 0.4 | – | – | 61.9 0.4 | 66.2 0.4 |
| ARC-C | accuracy | 54.9 1.5 | 59.9 1.4 | 59.8 1.4 | – | – | 59.0 1.4 | 60.0 1.4 |
| HellaSwag | accuracy | 70.8 0.5 | 73.1 0.4 | 73.4 0.4 | – | – | 72.8 0.4 | 73.2 0.4 |
| Dataset | Metric | Fixed-Depth | Dense-Ragged | |||
|---|---|---|---|---|---|---|
| depth / thr. | 2 | 3 | 4 | .3 | .5 | |
| Avg. depth | 2.00 | 3.00 | 4.00 | 2.21 | 3.03 | |
| Log-likelihood evaluation | ||||||
| MMLU | accuracy | 67.7 0.4 | 73.4 0.4 | 74.4 0.4 | – | – |
| ARC-C | accuracy | 63.5 1.4 | 65.4 1.4 | 66.6 1.4 | – | – |
| HellaSwag | accuracy | 76.8 0.4 | 77.9 0.4 | 78.4 0.4 | – | – |
| Dataset | Metric | Fixed-Depth | Dense-Ragged | ETE-Ragged | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| depth / thr. | 2 | 3 | 4 | .3 | .45 | .5 | .3 | .45 | .5 | |
| Avg. depth | 2.00 | 3.00 | 4.00 | 2.30 | 2.50 | 2.92 | 2.31 | 2.49 | 2.91 | |
| AIME24 | pass@1 | 24.2 5.4 | 47.9 7.3 | 50.4 7.1 | 45.8 7.1 | 47.9 7.0 | 46.2 6.9 | 40.8 6.7 | 45.8 7.0 | 46.2 6.8 |
| pass@4 | 48.5 7.6 | 67.3 7.9 | 71.7 7.7 | 67.1 8.2 | 69.7 7.7 | 68.6 7.9 | 64.9 7.8 | 68.0 7.9 | 69.5 7.6 | |
| AIME25 | pass@1 | 24.2 7.0 | 36.2 7.2 | 37.9 7.0 | 33.3 7.3 | 36.7 7.2 | 40.8 7.4 | 30.0 7.2 | 35.0 6.9 | 41.2 7.2 |
| pass@4 | 34.0 8.1 | 55.0 8.5 | 58.1 8.5 | 51.0 7.9 | 55.9 8.2 | 59.2 8.5 | 43.3 8.4 | 58.3 7.6 | 59.7 8.6 | |
| Dataset | Metric | Fixed-Depth | Dense-Ragged | ETE-Ragged | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| depth / thr. | 2 | 3 | 4 | .3 | .45 | .5 | .3 | .45 | .5 | |
| Avg. depth | 2.00 | 3.00 | 4.00 | 2.24 | 2.45 | 2.83 | 2.30 | 2.60 | 3.00 | |
| AIME24 | pass@1 | 40.4 6.7 | 62.1 7.0 | 65.8 6.8 | 56.7 7.3 | 57.9 7.3 | 61.7 7.2 | 55.0 7.1 | 59.6 7.2 | 61.7 6.9 |
| pass@4 | 64.4 7.7 | 79.9 6.9 | 82.0 6.9 | 76.0 7.3 | 75.2 7.4 | 78.7 6.8 | 77.3 6.4 | 76.4 7.3 | 80.4 6.9 | |
| AIME25 | pass@1 | 31.7 6.7 | 49.2 7.7 | 47.9 7.9 | 45.4 7.2 | 45.4 7.4 | 52.9 7.2 | 46.2 7.2 | 49.2 7.3 | 52.1 7.5 |
| pass@4 | 49.0 8.5 | 65.6 8.2 | 62.1 8.5 | 67.7 7.6 | 64.8 7.9 | 70.8 8.1 | 66.9 8.0 | 68.7 7.9 | 69.2 8.0 | |
| Dataset | Metric | Fixed-Depth | ETE-Ragged | ||
|---|---|---|---|---|---|
| depth / schedule | 3 | 4 | cos. / | ||
| Avg. depth (P / D / O) | 3.71 / 3.05 / 3.61 | 3.66 / 3.05 / 3.57 | |||
| Log-likelihood evaluation | |||||
| MMLU | accuracy | 66.7 0.4 | 67.5 0.4 | 66.3 0.4 | – |
| ARC-C | accuracy | 59.9 1.4 | 59.8 1.4 | 60.1 1.4 | – |
| HellaSwag | accuracy | 73.1 0.4 | 73.4 0.4 | 73.2 0.4 | – |
| Dataset | Metric | Fixed-Depth | ETE-Ragged | ||
|---|---|---|---|---|---|
| depth / schedule | 3 | 4 | cos. / | ||
| Avg. depth (P / D / O) | 3.63 / 3.03 / 3.51 | 4.00 / 3.45 / 3.88 | |||
| Log-likelihood evaluation | |||||
| MMLU | accuracy | 73.4 0.4 | 74.4 0.4 | 72.0 0.4 | – |
| ARC-C | accuracy | 65.4 1.4 | 66.6 1.4 | 65.4 1.4 | – |
| HellaSwag | accuracy | 77.9 0.4 | 78.4 0.4 | 78.1 0.4 | – |
| Dataset | Metric | Fixed-Depth | ETE-Ragged | ||
|---|---|---|---|---|---|
| depth / schedule | 3 | 4 | cos. / | ||
| Avg. depth (P / D / O) | 3.47 / 2.92 / 2.93 | 3.03 / 2.91 / 2.91 | |||
| AIME25 | pass@1 | 36.2 7.2 | 37.9 7.0 | 35.8 7.6 | 38.3 7.7 |
| pass@4 | 55.0 8.5 | 58.1 8.5 | 56.7 9.2 | 60.0 9.1 | |
| GPQA-Diamond | pass@1 | 38.3 2.3 | 38.4 2.3 | 39.6 2.5 | 37.1 2.5 |
| pass@4 | 67.9 2.7 | 67.5 2.7 | 70.7 3.2 | 66.2 3.4 | |
| Inference | P@1 | P@4 |
|---|---|---|
| Fixed@1 | 2.6 | 4.4 |
| Fixed@2 | 26.7 | 43.7 |
| Fixed@3 | 57.0 | 76.3 |
| Fixed@4 | 65.6 | 81.5 |