ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks
Organizations: Duke University · Tsinghua University
Abstract
Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response. ELF-REG-L achieves 55.96% pass@1 on GSM8K at 64 network function evaluations (NFE), and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. It outperforms the evaluated comparable-scale dLMs in pass@1 on GSM8K and code, and improves MATH-500 pass@1 from 10.55% for the ELF-L baseline to 13.39% with ELF-REG-L. Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory. At 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10, outperforming recent continuous dLMs of comparable scale.
Figures & tables
| Model | Param. | NFE / steps | MBPP-500 | MBPP-378 | HumanEval | HumanEval+ |
|---|---|---|---|---|---|---|
| Large dLMs/AR LLMs | ||||||
| LLaDA-8B-Base [ Nie et al., 2025 ] | 8.0B | n.r. | 38.80 | 52.60 | 35.40 | 30.50 |
| Dream-v0-Base-7B [ Ye et al., 2025 ] | 7.6B | n.r. | 55.40 | 71.50 | 56.70 | 50.00 |
| Qwen3-0.6B-Base/Qwen3-0.6B [ Yang et al., 2025 ] | 0.6B | – | – | – | 30.50 / 41.46 | – / 37.19 |
| Llama-3.2-1B/Llama-3.2-1B-Instruct [ Meta, 2024 ] | 1.2B | – | – | – | 18.90 / 33.50 | – |
| SmolLM2-1.7B/SmolLM2-1.7B-Instruct [ Allal et al., 2025 ] | 1.7B | – | – | – | 22.60 / 28.10 | – |
| Model | Param. | NFE / steps | Shots | pass@1 |
| GSM8K: large dLMs/AR LLMs | ||||
| LLaDA-8B-Base [ Nie et al., 2025 ] | 8.0B | 1,024 steps | 4 | 70.30 |
| Dream-v0-Base-7B [ Ye et al., 2025 ] | 7.6B | 256 steps | 8 | 77.20 |
| TESS 2 v0.1 (GSM8K fine-tuned) [ Tae et al., 2025 ] | 7B | 1,000 steps | 8 | 68.9 |
| Qwen3-0.6B-Base/Qwen3-0.6B [ Yang et al., 2025 ] | 0.6B | – | 4 / 0 | 59.59 / 79.20 |
| Llama-3.2-1B/Llama-3.2-1B-Instruct [ Meta, 2024 ] | 1.2B | – | 5 / 8 | 7.60 / 44.40 |
| Benchmark | Model | pass@2 | pass@4 | pass@8 | pass@16 |
|---|---|---|---|---|---|
| GSM8K | ELF-L baseline | 62.15 | 70.43 | 76.98 | 82.26 |
| GSM8K | ELF-REG-L (ours) | 65.18 | 72.57 | 78.64 | 83.70 |
| MATH-500 | ELF-L baseline | 15.65 | 23.64 | 33.36 | 44.00 |
| MATH-500 | ELF-REG-L (ours) | 20.75 | 29.95 | 40.41 | 52.20 |
| Baseline | REPA-only | REPA+REG | Stripped REPA+REG | REPA+OLT | |
| pass@1 (%) | 21.74 | 24.56 | 28.92 | 28.66 | 25.60 |
| NFE | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|
| Full-span | 7.25 | 15.75 | 25.27 | 30.95 | 34.19 |
| Prefix early-stop ( ) | 11.82 | 26.20 | 34.42 | 37.15 | 38.11 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | GSM8K | Code | MATH-500 | MMLU |
|---|---|---|---|---|
| Architecture | ||||
| Model size | ELF-B / ELF-L | ELF-L | ELF-L | ELF-B |
| Transformer blocks | 12 / 32 | 32 | 32 | 12 |
| Hidden width | 768 / 1,280 | 1,280 | 1,280 | 768 |
| Attention heads | 12 / 16 | 16 | 16 | 12 |
| SwiGLU hidden width | 2,048 / 3,413 | 3,413 | 3,413 | 2,048 |
| Task | Model | Epoch | Init. (B) | Task (B) | Total (B) | GPU-hours |
|---|---|---|---|---|---|---|
| GSM8K | ELF-B baseline | 18 | – | 10.044 | 10.044 | 170.8 |
| GSM8K | ELF-REG-B (ours) | 12 | – | 6.696 | 6.696 | 156.5 |
| GSM8K | ELF-L baseline | 15 | – | 8.370 | 8.370 | 347.9 |
| GSM8K | ELF-REG-L (ours) | 12 | – | 6.696 | 6.696 | 333.9 |
| Code | ELF-L baseline | 12 | – | 14.958 | 14.958 | 317.3 |
| Code | ELF-REG-L (ours) | 12 | – | 14.958 | 14.958 | 407.9 |
| GPU-hours | ||||
|---|---|---|---|---|
| Architecture | Accuracy (%) | Baseline | REPA+REG | Reduction (%) |
| ELF-B | 36.85 | 151.8 | 130.4 | 14.1 |
| ELF-L | 51.48 | 324.7 | 222.6 | 31.4 |
| Benchmark | Response cap (tokens) | Cap reached (%) |
|---|---|---|
| GSM8K | 2,048 | 0.76 |
| MATH-500 | 4,096 | 9.50 |
| MBPP-378 | 2,048 | 1.37 |
| HumanEval / HumanEval+ | 2,048 | 21.46 |
| Measurement | ELF-B baseline | ELF-REG-B (ours) |
|---|---|---|
| Numerical-answer agreement (%) | 80.27 | 87.11 |
| Complete-response agreement (%) | 1.46 | 8.61 |
| Incorrect correct, count (%) | 79 (1.50) | 70 (1.33) |
| Correct incorrect, count (%) | 78 (1.48) | 36 (0.68) |
| Accuracy at NFE 64 (%) | 36.71 | 37.64 |
| Accuracy at NFE 512 (%) | 36.73 | 38.29 |
| Quantity | Mean | Range |
|---|---|---|
| Prediction time | 0.0825 | 0.0755–0.0858 |
| Power SNR ( ) | 2.0272 | 1.6667–2.2043 |
| Condition | Median ( ) | Mean joint endpoint error |
|---|---|---|
| ELF-B baseline | 6.215 | 530.78 |
| ELF-REG-B (ours) | 0.946 | 490.80 |
| Model | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|
| ELF-L baseline | 27.64 | 43.20 | 48.58 | 50.82 | 51.85 |
| ELF-REG-L (ours) | 31.76 | 47.34 | 52.89 | 55.04 | 55.96 |
| ELF-B baseline | 10.24 | 25.88 | 33.85 | 36.81 | 37.71 |
| ELF-REG-B (ours) | 11.82 | 26.20 | 34.42 | 37.15 | 38.11 |
| FMLM+ (Init) [ Agarwal et al., 2026 ] | 5.10 | 15.10 | 26.10 | 31.80 | 33.60 |
| DBTM [ Tang & Wang, 2026 ] | 5.70 | 10.10 | 14.30 | 16.80 | 16.20 |
| Model | NFE / steps | MBPP-500 | MBPP-378 | HumanEval | HumanEval+ |
|---|---|---|---|---|---|
| ELF-L baseline | 4 NFE | 4.20 | 8.10 | 5.68 | 5.30 |
| ELF-L baseline | 8 NFE | 9.68 | 16.19 | 10.44 | 9.76 |
| ELF-L baseline | 16 NFE | 12.89 | 20.42 | 14.10 | 13.38 |
| ELF-L baseline | 32 NFE | 15.00 | 23.84 | 17.19 | 16.27 |
| ELF-L baseline | 64 NFE | 15.95 | 25.69 | 18.33 | 17.49 |
| ELF-L baseline | 128 NFE | 17.19 | 26.75 | 19.66 | 18.75 |
| Model | NFE / steps | MBPP-500 | MBPP-378 | HumanEval | HumanEval+ |
|---|---|---|---|---|---|
| ELF-L baseline | 4 NFE | 16.30 | 26.84 | 16.39 | 15.25 |
| ELF-L baseline | 8 NFE | 28.13 | 42.86 | 30.43 | 28.53 |
| ELF-L baseline | 16 NFE | 34.75 | 47.57 | 35.67 | 34.15 |
| ELF-L baseline | 32 NFE | 37.57 | 50.43 | 42.93 | 40.09 |
| ELF-L baseline | 64 NFE | 38.41 | 52.88 | 45.04 | 42.28 |
| ELF-L baseline | 128 NFE | 40.50 | 52.90 | 45.93 | 42.84 |
| Model | NFE | pass@1 | pass@2 | pass@4 | pass@8 | pass@10 | pass@16 |
|---|---|---|---|---|---|---|---|
| MBPP-500 | |||||||
| ELF-L baseline | 128 | 17.19 | 24.10 | 31.28 | 38.34 | 40.50 | 44.80 |
| ELF-REG-L (ours) | 128 | 19.10 | 25.91 | 32.41 | 38.18 | 39.92 | 43.40 |
| PlaidQ [ Peng et al., 2026b ] | 129 | 12.19 | – | – | – | 21.44 | – |
| PlaidQ+CFG [ Peng et al., 2026b ] | 257 | 15.53 | – | – | – | 32.48 | – |
| MBPP-378 | |||||||
| Model | 4 | 8 | 16 | 32 | 64 | 128 |
|---|---|---|---|---|---|---|
| ELF-L baseline | 5.47 | 7.58 | 8.56 | 9.65 | 9.65 | 10.55 |
| ELF-REG-L (ours) | 7.04 | 9.76 | 11.53 | 12.40 | 13.31 | 13.39 |
| Model | pass@1 | pass@2 | pass@4 | pass@8 | pass@10 | pass@16 |
|---|---|---|---|---|---|---|
| ELF-L baseline | 9.65 | 15.65 | 23.64 | 33.36 | 36.73 | 44.00 |
| ELF-REG-L (ours) | 13.31 | 20.75 | 29.95 | 40.41 | 44.04 | 52.20 |
| Model | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|
| ELF-B baseline | 28.67 | 44.00 | 45.23 | 45.32 | 45.34 |
| ELF-REG-B (ours) | 30.27 | 45.60 | 46.66 | 46.69 | 46.71 |
| Setting | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|
| FMLM+ (Init) [ Agarwal et al., 2026 ] | 5.10 | 15.10 | 26.10 | 31.80 | 33.60 |
| FMLM+ (Init) (our reproduction) [ Agarwal et al., 2026 ] | 6.05 | 15.85 | 25.38 | 31.37 | 33.36 |
| FMLM+ (Init) early-stop, master NFE 64 [ Agarwal et al., 2026 ] | 2.31 | 6.25 | 16.83 | 29.64 | 33.36 |
| ELF-REG-L (ours) | ||||||
|---|---|---|---|---|---|---|
| NFE | 4 | 8 | 16 | 32 | 64 | 128 |
| pass@10 | 31.48 | 46.39 | 49.99 | 52.78 | 54.98 | 55.05 |
| Hidden-state change ( ) | ||||||
| Block | Condition | |||||
| 4 | baseline | 1.562 | 0.120 | 0.123 | 0.003 | 0.016 |
| 4 | REPA+REG shift-0 | 2.422 | 0.426 | 0.444 | 0.019 | 0.019 |
| 4 | REPA+REG shift-1 | 1.867 | 0.407 | 0.449 | 0.042 | 0.045 |
| 12 | baseline | 0.176 | 0.068 | 0.095 | 0.027 | 0.117 |
| Condition | Validation pass@1 (%) | Test pass@1 (%) |
|---|---|---|
| ELF-B baseline | 21.58 | 17.89 |
| ELF-B baseline, Qwen3-1.7B-Base encoder | 10.16 | 10.39 |
| ELF-REG-B (ours) (default) | 32.32 | 25.02 |
| Text REPA during denoising: | 31.05 | 25.17 |
| REG target: mean pooling | 27.05 | 23.65 |
| REG only | 25.39 | 17.97 |
| Condition | Validation pass@1 (%) | Test pass@1 (%) |
|---|---|---|
| ELF-L baseline | 33.50 | 27.37 |
| ELF-REG-L (ours) (default) | 44.04 | 34.04 |
| Alignment depth: 6 | 42.09 | 33.13 |
| Alignment depth: 10 | 41.21 | 33.28 |
| Alignment depth: 11 | 40.04 | 33.13 |
| / NFE | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|
| 1 | 7.25 | 15.75 | 25.27 | 30.95 | 34.19 |
| 2 | 10.60 | 23.22 | 30.89 | 34.08 | 35.84 |
| 4 | 11.62 | 25.75 | 33.43 | 35.69 | 37.62 |
| 8 | 11.82 | 26.20 | 34.42 | 37.15 | 38.11 |
| 16 | 11.22 | 25.82 | 35.30 | 37.55 | 38.32 |
| NFE | ELF-B baseline | ELF-REG-B (ours) |
|---|---|---|
| 4 | 13.87 | 17.86 |
| 8 | 39.59 | 44.62 |
| 16 | 65.93 | 71.34 |
| 32 | 81.95 | 85.06 |
| 64 | 100.00 | 100.00 |
| A. Language generation | |||
| Method | Generation state | Response generation | Training or sampling approach |
| LangFlow [ Chen et al., 2026 ] | Token embeddings | Parallel sequence denoising | Flow matching with a learned noise schedule. |
| LDLM [ Meshchaninov et al., 2026 ] | Learned contextual latents | Parallel denoising, then token decoding | Joint training of the latent encoder, denoiser, and decoder. |
| ELF [ Hu et al., 2026 ] | Contextual encoder representations | Parallel denoising; final token decoding | Flow matching with a shared denoiser and decoder. |
| S-FLM [ Deschenaux & Gulcehre, 2026 ] | Hyperspherical token embeddings | Parallel sequence denoising | Riemannian flow matching on learned embeddings. |
| FMLM [ Lee et al., 2026 ] | Noisy one-hot token vectors | Parallel sequence generation | Distillation of a continuous flow into a flow map. |