Reasoning with Continuous Latent Diffusion
Organizations: Duke University Department of Electrical and Computer Engineering
Abstract
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce the Continuous Embedding Diffusion Reasoner (CEDR), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT CEDR-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/CEDR.
Figures & tables
| Task / model | Frozen-Qwen flow epochs | Prompt-MSE epochs | Joint epochs (flow endpoint) | Reported NFT update |
|---|---|---|---|---|
| GSM8K / CEDR-B | 12 | 60 | 6 (18) | 300 |
| GSM8K / CEDR-L | 12 | 30 | 9 (21) | 500 |
| MATH / CEDR-L | 12 | 10 | 9 (21) | 600 |
| OpenCodeInstruct / CEDR-L | 12 | 10 | 1 (13) | 100 |
| GSM8K | MATH500 | |
|---|---|---|
| 52.96 ( ) | 15.50 ( ) | |
| 57.51 ( ) | 17.40 ( ) | |
| 58.07 ( ) | 17.60 ( ) | |
| 59.25 ( ) | 17.90 ( ) |
| Model | Embedding Params (M) | Other Params (M) | NFE | GSM8K | MATH500 |
|---|---|---|---|---|---|
| Autoregressive | |||||
| Qwen3-0.6B-Base ( Yang et al., 2025 ) | 156 | 440 | — | 59.59 | — |
| Qwen3-0.6B, thinking ( Yang et al., 2025 ) | 156 | 440 | — | — | 77.60 |
| Discrete diffusion / blockwise diffusion | |||||
| LLaDA-8B-Base ( Nie et al., 2025 ) | 1,040 | 6,980 | 1,024 | 70.30 | — |
| Dream-v0-Base-7B ( Ye et al., 2025 ) | 1,090 | 6,530 | 256 | 77.20 | — |
| GSM8K ( 100M) | GSM8K ( 650M) | MATH500 ( 650M) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method / NFE | 8 | 16 | 32 | 64 | 8 | 16 | 32 | 64 | 8 | 16 | 32 | 64 |
| FMLM+ Init | 15.10 | 26.10 | 31.80 | 33.60 | — | — | — | — | — | — | — | — |
| ELF-REG | ∗ 26.20 | ∗ 34.42 | ∗ 37.15 | ∗ 38.11 | ∗ 47.34 | ∗ 52.89 | ∗ 55.04 | ∗ 55.96 | ∗ 9.76 | ∗ 11.53 | ∗ 12.40 | ∗ 13.31 |
| CEDR, pre-NFT | 26.40 | 34.48 | 37.80 | 39.17 | 49.41 | 55.33 | 57.88 | 58.67 | 15.23 | 17.93 | 19.65 | 21.18 |
| CEDR, post-NFT | 30.08 | 37.77 | 41.13 | 42.16 | 54.84 | 60.55 | 62.47 | 63.74 | 17.45 | 20.93 | 22.60 | 24.68 |
| Model | NFE | Pass@ (%) | ||
|---|---|---|---|---|
| ELF-REG-L ( Li et al., 2026 ) | 63 | 8 | 504 | 40.41 |
| ELF-REG-L ( Li et al., 2026 ) | 63 | 10 | 630 | 44.04 |
| CEDR-L, pre-NFT | 64 | 8 | 512 | 47.20 |
| CEDR-L, post-NFT | 64 | 8 | 512 | 52.60 |
| Model / setting | Denoising NFE | HE | HE+ | MBPP-378 base | MBPP+ full |
|---|---|---|---|---|---|
| Standard scoring: no function-name repair | |||||
| PlaidQ ( Peng et al., 2026b ) | 128 | 19.02 | 18.05 | 22.18 | — |
| PlaidQ+CFG ( Peng et al., 2026b ) | 256 | 22.04 | 19.70 | 24.58 | — |
| CEDR-L / frozen Qwen | 128 | 33.69 | 30.87 | 38.72 | 33.99 |
| CEDR-L / prompt MSE 10 | 128 | 27.97 | 25.30 | 19.61 | 16.96 |
| CEDR-L, pre-NFT (ours) | 128 | 29.57 | 26.91 | 18.45 | 16.40 |
| Model | NFE | HE | HE+ | MBPP-378 base | MBPP+ full | |
| Standard scoring: no function-name repair | ||||||
| PlaidQ ( Peng et al., 2026b ) | 128 | 10 | 25.72 | 24.03 | 34.33 | — |
| PlaidQ+CFG ( Peng et al., 2026b ) | 256 | 10 | 39.87 | 36.74 | 44.79 | — |
| CEDR-L, pre-NFT | 128 | 8 | 54.27 | 47.56 | 39.95 | 35.98 |
| CEDR-L, post-NFT | 128 | 8 | 54.27 | 49.39 | 43.39 | 37.83 |
| Single-function name aliasing | ||||||
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | CEDR-B | CEDR-L |
|---|---|---|
| ELF backbone: blocks / width / heads | 12 / 768 / 12 | 32 / 1,280 / 16 |
| ELF input bottleneck | 512 | 512 |
| Prompt Transformer: blocks / width / heads | 6 / 768 / 12 | 6 / 1,280 / 16 |
| Prompt frontend | ||
| Prompt output width | 1,024 | 1,024 |
| Frozen Qwen token lookup | ||
| Source | Retained subset | Training rows |
|---|---|---|
| OpenMathInstruct-2 | gsm8k | 292,485 |
| augmented_gsm8k | 1,841,419 | |
| Orca-Math | Eligible train rows | 147,464 |
| MetaMathQA | GSM types | 152,962 |
| Total | 2,434,330 |
| Epoch | Prompt | Correct 42 | Correct 123 | Mean (SD) | 95% CI |
| GSM8K, CEDR-B; | |||||
| 4 | Qwen | 92 | 86 | 6.75 (0.32) | |
| 8 | Qwen | 401 | 424 | 31.27 (1.23) | |
| 12 | Qwen | 514 | 503 | 38.55 (0.59) | |
| 12 | MSE swap | 398 | 381 | 29.53 (0.91) | |
| 16 | Joint | 482 | 462 | 35.78 (1.07) | |
| Probe input | Token accuracy (%) | CE (nats/token) |
|---|---|---|
| Fixed layer-32 PCA | 99.24 [99.19, 99.28] | 0.0548 [0.0502, 0.0597] |
| Epoch | Learned 16/24/32 | PCA 16/24/32 | PCA layer 32 | Epoch-0 joint prompt training |
|---|---|---|---|---|
| 1 | 3.75 (0.05) | 4.85 (0.75) | 2.65 (0.11) | 2.46 (0.05) |
| 2 | 12.40 (0.27) | 11.18 (0.16) | 5.50 (0.70) | 3.98 (1.02) |
| 3 | 20.17 (0.54) | 16.38 (0.32) | 8.64 (0.43) | 3.68 (0.16) |
| 4 | 26.61 (0.86) | 21.49 (0.80) | 7.47 (0.05) | 7.77 (1.02) |
| 5 | 29.95 (0.21) | 25.85 (0.21) | 9.82 (0.16) | 7.54 (1.45) |
| Validation | Test | |||
|---|---|---|---|---|
| Clock powers | GSM8K | MATH | GSM8K | MATH500 |
| 65.00 (0.99) | 15.40 (0.28) | 52.96 (1.13) | 15.50 (1.27) | |
| 72.70 (0.85) | 17.80 (0.57) | 57.51 (0.80) | 17.40 (1.41) | |
| 72.80 (1.13) | 18.30 (0.14) | 58.07 (0.64) | 17.60 (1.98) | |
| 72.15 (0.35) | 18.70 (0.14) | 59.25 (0.16) | 17.90 (1.27) | |
| Layer | Actual group [95% CI] | Random mean [min, max] | |
|---|---|---|---|
| First prediction | |||
| 16 | 0 | ||
| 16 | 0.8 | ||
| 16 | 1.6 | ||
| 16 | 2 | ||
| 16 | 3 | ||
| Model / stage | 8 NFE | 16 NFE | 32 NFE | 64 NFE |
|---|---|---|---|---|
| Sampling seeds | 64 | 32 | 16 | 8 |
| GSM8K, CEDR-B, SCCFG 2 | ||||
| Pre-NFT | ||||
| Post-NFT | ||||
| GSM8K, CEDR-L, SCCFG 2 | ||||
| Pre-NFT | ||||
| Oracle pass@ | Majority vote | |||||
| NFE | Pre-NFT | Post-NFT | Pre-NFT | Post-NFT | ||
| 8 | 1 | 8 | 49.41 (1.04) | 54.84 (1.09) | 49.41 (1.04) | 54.84 (1.09) |
| 8 | 2 | 16 | 60.21 (1.08) | 63.89 (1.11) | 49.41 (1.04) | 54.84 (1.09) |
| 8 | 4 | 32 | 68.84 (1.05) | 70.85 (1.08) | 56.11 (1.15) | 60.04 (1.17) |
| 8 | 8 | 64 | 76.01 (0.99) | 76.53 (1.02) | 59.92 (1.18) | 62.69 (1.20) |
| 8 | 16 | 128 | 81.85 (0.92) | 81.34 (0.95) | 62.05 (1.21) | 64.07 (1.24) |
| Oracle pass@ | Majority vote | |||||
| NFE | Pre-NFT | Post-NFT | Pre-NFT | Post-NFT | ||
| 8 | 1 | 8 | 13.88 (0.99) | 16.31 (1.11) | 13.88 (0.99) | 16.31 (1.11) |
| 8 | 2 | 16 | 21.10 (1.29) | 23.98 (1.40) | 13.88 (0.99) | 16.31 (1.11) |
| 8 | 4 | 32 | 29.74 (1.55) | 32.65 (1.63) | 16.87 (1.24) | 19.64 (1.37) |
| 8 | 8 | 64 | 39.29 (1.73) | 41.82 (1.79) | 20.00 (1.44) | 22.68 (1.56) |
| 8 | 16 | 128 | 49.27 (1.86) | 50.94 (1.90) | 22.23 (1.58) | 24.59 (1.69) |
| Oracle pass@ | Majority vote | |||||
| NFE | Pre-NFT | Post-NFT | Pre-NFT | Post-NFT | ||
| 8 | 1 | 8 | 26.40 (0.86) | 30.08 (0.95) | 26.40 (0.86) | 30.08 (0.95) |
| 8 | 2 | 16 | 36.10 (1.02) | 39.21 (1.08) | 26.40 (0.86) | 30.08 (0.95) |
| 8 | 4 | 32 | 45.58 (1.11) | 47.77 (1.14) | 31.20 (1.02) | 34.14 (1.09) |
| 8 | 8 | 64 | 54.40 (1.15) | 55.75 (1.16) | 34.55 (1.11) | 36.50 (1.16) |
| 8 | 16 | 128 | 62.54 (1.15) | 63.20 (1.15) | 36.49 (1.18) | 37.77 (1.22) |
| Checkpoint / conditioner | Seeds | HE | HE+ | MBPP-378 base | MBPP+ |
|---|---|---|---|---|---|
| Epoch 4 / Qwen | 2 | 23.17 (3.45) | 21.04 (2.16) | 26.85 (0.19) | 23.68 (0.19) |
| Epoch 8 / Qwen | 2 | 27.74 (6.47) | 25.61 (5.17) | 34.92 (2.24) | 30.42 (1.87) |
| Epoch 12 / Qwen | 4 | 27.90 (1.75) | 25.61 (2.49) | 37.76 (0.70) | 32.34 (0.79) |
| Epoch 12 / prompt MSE 1 | 4 | 19.82 (2.08) | 18.60 (2.31) | 8.93 (1.02) | 7.61 (1.23) |
| Epoch 12 / prompt MSE 10 | 4 | 26.37 (1.44) | 24.70 (2.08) | 17.59 (2.61) | 15.15 (2.16) |
| Stage | Benchmark | pass@1 | pass@2 | pass@4 | pass@8 | pass@1 seed SD |
|---|---|---|---|---|---|---|
| Frozen Qwen | HumanEval | 33.69 (2.78) | 44.66 (3.21) | 54.13 (3.46) | 62.20 (3.75) | 2.11 |
| Frozen Qwen | HumanEval+ | 30.87 (2.70) | 41.53 (3.18) | 50.82 (3.45) | 59.15 (3.77) | 2.09 |
| Frozen Qwen | MBPP-378 base | 38.72 (2.01) | 49.17 (2.21) | 58.26 (2.32) | 66.40 (2.48) | 1.37 |
| Frozen Qwen | MBPP+ | 33.99 (2.00) | 43.26 (2.25) | 51.46 (2.40) | 58.73 (2.60) | 1.62 |
| Prompt MSE | HumanEval | 27.97 (2.84) | 35.98 (3.27) | 42.87 (3.58) | 48.78 (3.90) | 1.44 |
| Prompt MSE | HumanEval+ | 25.30 (2.81) | 32.19 (3.25) | 37.65 (3.55) | 42.07 (3.85) | 1.84 |
| Stage | Benchmark | pass@1 | pass@2 | pass@4 | pass@8 | pass@1 seed SD |
|---|---|---|---|---|---|---|
| Frozen Qwen | HumanEval | 33.69 (2.78) | 44.66 (3.21) | 54.13 (3.46) | 62.20 (3.75) | 2.11 |
| Frozen Qwen | HumanEval+ | 30.87 (2.70) | 41.53 (3.18) | 50.82 (3.45) | 59.15 (3.77) | 2.09 |
| Frozen Qwen | MBPP-378 base | 43.45 (2.06) | 53.97 (2.22) | 62.90 (2.28) | 70.63 (2.39) | 1.21 |
| Frozen Qwen | MBPP+ | 37.93 (2.07) | 47.21 (2.28) | 55.07 (2.40) | 61.90 (2.55) | 1.43 |
| Prompt MSE | HumanEval | 28.05 (2.84) | 36.11 (3.27) | 43.05 (3.59) | 48.78 (3.90) | 1.38 |
| Prompt MSE | HumanEval+ | 25.38 (2.81) | 32.32 (3.25) | 37.82 (3.56) | 42.07 (3.85) | 1.78 |
| Scoring | Benchmark | pass@1 [95% CI] | pass@8 [95% CI] |
|---|---|---|---|
| Standard | HumanEval | ||
| Standard | HumanEval+ | ||
| Standard | MBPP-378 base | ||
| Standard | MBPP+ | ||
| Alias | HumanEval | ||
| Alias | HumanEval+ |
| Scoring | Benchmark | pass@1 [95% CI] | pass@8 [95% CI] |
|---|---|---|---|
| Standard | HumanEval | ||
| Standard | HumanEval+ | ||
| Standard | MBPP-378 base | ||
| Standard | MBPP+ | ||
| Alias | HumanEval | ||
| Alias | HumanEval+ |