Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models
Organizations: NVIDIA · Cornell University
Abstract
Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.
Figures & tables
| NFEs/seq. ( ) | ROUGE ( ) | |||||
| Mean | 1 | 2 | L | |||
| Autoregressive | 73.1 | 35.9 | 14.8 | 25.0 | ||
| Discrete Diffusion Baselines | ||||||
| MDLM | (full seq.) | 180.0 | 39.5 | 17.1 | 26.1 | |
| ( Sahoo et al., 2024a ) | ( ) | 128.0 | 40.1 | 18.1 | 27.2 | |
| ( ) | 96.0 | 40.0 | 18.0 | 27.1 | ||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Data | |
|---|---|
| Tokenizer / vocabulary | gpt2 / |
| Context length | |
| Packing | [EOS] per document; [BOS] / [EOS] per packed block |
| Validation split | Last documents |
| Architecture | |
| Backbone | DiT ( Peebles et al., 2023 ) with adaptive layer norm |
| Data | |
|---|---|
| Tokenizer / vocabulary | HuggingFaceTB/SmolLM-135M / |
| Context length | (right-padded, overlong examples filtered) |
| Packing | None; one example per sequence |
| Prompt / target | [BOS] question \n / code [EOS] |
| Validation split | deterministic holdout of TinyGSM train |
| Architecture | |
| Data | |
|---|---|
| Tokenizer / vocabulary | Qwen/Qwen3-0.6B-Base / |
| Context length | ( source target, right-padded) |
| Target prefix | "Summary: " |
| Supervision | Target region only; padding neither noised nor scored |
| Architecture | |
| Backbone | Qwen3-style decoder (architecture only; trained from scratch) |
| NFEs/seq. ( ) | Acc. % ( ) | NFEs/seq. ( ) | Acc. % ( ) | |||
| Mean | -shot, pass@ | Mean | -shot, pass@ | |||
| B-ClockDLM | ||||||
| Standard | 190.1 | 54.3 | 220.1 | 49.1 | ||
| + Cache Grab-Conf. | ( ) | 134.8 | 51.6 | 182.3 | 48.5 | |
| ( ) | 117.5 | 53.8 | 161.8 | 47.2 | ||
| Diffusion loss only | Joint AR diffusion loss | ||||
|---|---|---|---|---|---|
| NFEs/seq. ( ) | Acc. % ( ) | NFEs/seq. ( ) | Acc. % ( ) | ||
| 4 | 161.5 | 65.7 | 160.4 | 66.6 | 0.9 |
| 8 | 165.0 | 63.8 | 164.5 | 67.2 | 3.4 |
| 16 | 172.6 | 60.0 | 173.7 | 63.8 | 3.8 |
| 32 | 188.1 | 59.6 | 187.4 | 63.5 | 3.9 |
| 64 | 219.1 | 55.6 | 219.5 | 61.3 | 5.7 |
| Dataset | License | URL |
|---|---|---|
| OpenWebText ( Gokaslan et al., 2019 ) | CC0 1.0 | https://huggingface.co/datasets/Skylion007/openwebtext |
| TinyGSM ( Liu et al., 2023 ) | MIT | https://huggingface.co/datasets/TinyGSM/TinyGSM |
| GSM8K ( Cobbe et al., 2021 ) | MIT | https://github.com/openai/grade-school-math |
| CNN/DailyMail ( Hermann et al., 2015 ; See et al., 2017 ) | Apache 2.0 | https://huggingface.co/datasets/abisee/cnn_dailymail |
| Asset | Use | License | URL |
|---|---|---|---|
| GPT-2 ( Radford et al., 2019 ) | Tokenizer | MIT | https://huggingface.co/openai-community/gpt2 |
| SmolLM-135M ( Allal et al., 2024 ) | Tokenizer | Apache 2.0 | https://huggingface.co/HuggingFaceTB/SmolLM-135M |
| Qwen3-0.6B-Base ( Team, 2025 ) | Tokenizer, architecture | Apache 2.0 | https://huggingface.co/Qwen/Qwen3-0.6B-Base |
| ModernBERT ( Warner et al., 2025 ) | MAUVE evaluation | Apache 2.0 | https://huggingface.co/answerdotai/ModernBERT-large |
| -FLM ( Deschenaux et al., 2026 ) | Baseline models, code | Apache 2.0 | https://huggingface.co/jdeschena/s-flm |
| FMLM ( Agarwal et al., 2026 ) | Baseline models, code | Apache 2.0 | https://github.com/MananAg007/posterior-refinement |
| Library | License |
|---|---|
| HuggingFace ( Wolf et al., 2019 ) | Apache 2.0 |
| Hydra ( Yadan, 2019 ) | MIT |
| Language Model Evaluation Harness ( Gao et al., 2021 ) | MIT |
| Matplotlib ( Hunter, 2007 ) | Matplotlib License |
| MAUVE ( Pillutla, 2021 ) | GNU General Public License, Version 3 |
| NumPy ( Harris et al., 2020 ) | BSD 3-Clause |