Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with representation alignment and entanglement (REPA+REG), where a frozen AR teacher supervises intermediate denoiser features and supplies a global representation that is jointly denoised with the response. ELF-REG-L achieves 55.96% pass@1 on GSM8K at 64 network function evaluations (NFE), and 13.39% on MATH-500 and 22.56% on HumanEval at 128 NFE. It outperforms the evaluated comparable-scale dLMs in pass@1 on GSM8K and code, and improves MATH-500 pass@1 from 10.55% for the ELF-L baseline to 13.39% with ELF-REG-L. Without few-step training, the same task-specific checkpoints support strong low-NFE performance through early-stop, which decodes an intermediate clean prediction without completing the denoising trajectory. At 16 NFE, ELF-REG-L reaches 41.21% HumanEval pass@10, outperforming recent continuous dLMs of comparable scale.
Figures & tables
Figure 1: ELF-REG improves generation with fewer task-training tokens. (a) A frozen AR teacher supervises intermediate features through REPA and a global REG token denoised jointly with the response. (b,c) ELF-REG-L and ELF-REG-B with early-stop ( ρ=8 ) show superior performance across NFE, compared with PlaidQ [ Peng et al., 2026b ] , FMLM+ [ Agarwal et al., 2026 ] , S-FLM [ Deschenaux & Gulcehre, 2026 ] , and DBTM [ Tang & Wang, 2026 ] . Code uses pass@10 on MBPP-378; GSM8K uses mean pass@1. (d) Non-padding training token budgets. Appendix B gives more details regarding benchmark mappings, model sizes, NFE accounting, budget estimates, and evaluation differences.
Model
Param.
NFE / steps
MBPP-500
MBPP-378
HumanEval
HumanEval+
Large dLMs/AR LLMs
LLaDA-8B-Base [ Nie et al., 2025 ]
8.0B
n.r.
38.80
52.60
35.40
30.50
Dream-v0-Base-7B [ Ye et al., 2025 ]
7.6B
n.r.
55.40
71.50
56.70
50.00
Qwen3-0.6B-Base/Qwen3-0.6B [ Yang et al., 2025 ]
0.6B
–
–
–
30.50 / 41.46
– / 37.19
Llama-3.2-1B/Llama-3.2-1B-Instruct [ Meta, 2024 ]
1.2B
–
–
–
18.90 / 33.50
–
SmolLM2-1.7B/SmolLM2-1.7B-Instruct [ Allal et al., 2025 ]
1.7B
–
–
–
22.60 / 28.10
–
Table 1: Zero-shot code generation, pass@1 (%). Our evaluations report means, with standard deviation across 16 seeds in parentheses. Slashes separate base/instruct names and scores; Qwen3-0.6B results use non-thinking mode. Bold marks the best comparable-scale dLM mean per benchmark. Both MBPP columns use base tests; HumanEval+ uses extended tests. Parameter counts follow Table 2 . Budgets specify NFE or sampling steps; n.r. means unreported. Sources and protocols appear in Appendix B .
Table 2: Mathematical reasoning, pass@1 (%). Our evaluations report mean with standard deviation across 16 seeds in parentheses. Slashes separate base/instruct models, scores, and differing shot counts. Qwen3-0.6B MATH-500 accuracy uses thinking mode. Bold marks the best comparable-scale dLM performance per benchmark. We show parameter count as non-decoder(+decoder) for ELF, including REPA+REG components. Budgets specify NFE or sampling steps; shots count demonstrations, and n.r. means unreported. Sources and protocols appear in Appendix B .
Benchmark
Model
pass@2
pass@4
pass@8
pass@16
GSM8K
ELF-L baseline
62.15
70.43
76.98
82.26
GSM8K
ELF-REG-L (ours)
65.18
72.57
78.64
83.70
MATH-500
ELF-L baseline
15.65
23.64
33.36
44.00
MATH-500
ELF-REG-L (ours)
20.75
29.95
40.41
52.20
Table 3: Mathematical reasoning, pass@ k (%), at NFE 64. ELF-REG-L improves ELF-L baseline across the sample budgets on both benchmarks. Results are shown across 16 seeds. Bold marks the best value per benchmark and sample budget.
Baseline
REPA-only
REPA+REG
Stripped REPA+REG
REPA+OLT
pass@1 (%)
21.74
24.56
28.92
28.66
25.60
Table 4: REPA+REG training benefits largely survive removal of REG at inference. At epoch 6 on GSM8K, REPA-only improves over Baseline, and REPA+REG improves further. Stripped REPA+REG retains most of this additional gain using the same trained weights as REPA+REG, except the REG component which is removed. REPA+OLT adds one learnable token (OLT) to REPA and does not reproduce the REG benefit. All settings use ELF-B. The REPA+REG setting corresponds to ELF-REG-B (ours).
NFE
4
8
16
32
64
Full-span
7.25
15.75
25.27
30.95
34.19
Prefix early-stop ( ρ=8 )
11.82
26.20
34.42
37.15
38.11
Table 5: GSM8K pass@1 (%) at matched NFE, ELF-REG-B (ours) at epoch 12, averaged over sixteen generation seeds.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
GSM8K
Code
MATH-500
MMLU
Architecture
Model size
ELF-B / ELF-L
ELF-L
ELF-L
ELF-B
Transformer blocks
12 / 32
32
32
12
Hidden width
768 / 1,280
1,280
1,280
768
Attention heads
12 / 16
16
16
12
SwiGLU hidden width
2,048 / 3,413
3,413
3,413
2,048
Appendix
Table 6: Default hyperparameters for the headline ELF configurations. GSM8K pairs give ELF-B / ELF-L when the sizes differ; other slashes separate the quantities named in the row. Parameters are non-decoder (+decoder), in millions. REPA+REG denotes ELF-REG-B (ours) or ELF-REG-L (ours), according to model size; the baseline omits both objectives. Training epochs identify the evaluated checkpoints. MATH-500 epochs count task training after GSM8K initialization (Appendix B.2 ). Shared entries span the task columns.
Task
Model
Epoch
Init. (B)
Task (B)
Total (B)
GPU-hours
GSM8K
ELF-B baseline
18
–
10.044
10.044
170.8
GSM8K
ELF-REG-B (ours)
12
–
6.696
6.696
156.5
GSM8K
ELF-L baseline
15
–
8.370
8.370
347.9
GSM8K
ELF-REG-L (ours)
12
–
6.696
6.696
333.9
Code
ELF-L baseline
12
–
14.958
14.958
317.3
Code
ELF-REG-L (ours)
12
–
14.958
14.958
407.9
Appendix
Table 7: Training epochs, non-padding task-training tokens, and GPU-hours on H200 GPUs for the headline checkpoints. Init. counts inherited task training; Task counts training on the current corpus; Total sums them. Token counts are epoch-based estimates in billions. MATH-500 GPU-hours are given including / excluding GSM8K initialization.
GPU-hours
Architecture
Accuracy (%)
Baseline
REPA+REG
Reduction (%)
ELF-B
36.85
151.8
130.4
14.1
ELF-L
51.48
324.7
222.6
31.4
Appendix
Table 8: Training GPU-hours on H200 GPUs to reach the GSM8K accuracy thresholds in Figure 3 . Reduction is relative to the baseline of the same architecture. REPA+REG denotes ELF-REG-B (ours) or ELF-REG-L (ours), according to architecture.
Benchmark
Response cap (tokens)
Cap reached (%)
GSM8K
2,048
0.76
MATH-500
4,096
9.50
MBPP-378
2,048
1.37
HumanEval / HumanEval+
2,048
21.46
Appendix
Table 9: Response-token caps and the percentage of Qwen3-0.6B-Base [ Yang et al., 2025 ] generations that reach them in our evaluations. Each question has sixteen samples. HumanEval and HumanEval+ share the same generations.
Measurement
ELF-B baseline
ELF-REG-B (ours)
Numerical-answer agreement (%)
80.27
87.11
Complete-response agreement (%)
1.46
8.61
Incorrect → correct, count (%)
79 (1.50)
70 (1.33)
Correct → incorrect, count (%)
78 (1.48)
36 (0.68)
Accuracy at NFE 64 (%)
36.71
37.64
Accuracy at NFE 512 (%)
36.73
38.29
Appendix
Table 10: Numerical-answer agreement and correctness transitions from early-stop NFE 64 to the endpoint at NFE 512 on GSM8K. Each early-stop prediction is paired with the endpoint of its own master-NFE-512 trajectory. ELF-B baseline uses epoch 18 and ELF-REG-B epoch 12. Each column covers 5,276 question–seed pairs: all 1,319 test questions and four generation seeds. Most numerical answers agree despite low complete-response agreement; correctness gains and losses nearly cancel for ELF-B baseline.
Quantity
Mean
Range
Prediction time t
0.0825
0.0755–0.0858
Power SNR ( ×10−3 )
2.0272
1.6667–2.2043
Appendix
Table 11: Prediction time and nominal training-interpolant power SNR at the NFE-64 exit ( ρ=8 ) in Figure 2 . Means and ranges are over four saved logit-normal schedules, shared by ELF-B baseline and ELF-REG-B (ours). SNR is computed separately for each schedule before averaging.
Condition
Median A^ ( ×105 )
Mean joint endpoint error
ELF-B baseline
6.215
530.78
ELF-REG-B (ours)
0.946
490.80
Appendix
Table 12: Remaining joint acceleration and joint endpoint error at t=5/64 on GSM8K, using 512 Euler updates with uniform time sampling. Each early-stop prediction is compared with its own trajectory endpoint. Acceleration is the median over all 1,319 questions within each generation seed; endpoint error is the mean Euclidean norm. Both statistics then average four generation seeds. ELF-B baseline uses epoch 18 and ELF-REG-B epoch 12.
Model
4
8
16
32
64
ELF-L baseline
27.64
43.20
48.58
50.82
51.85
ELF-REG-L (ours)
31.76
47.34
52.89
55.04
55.96
ELF-B baseline
10.24
25.88
33.85
36.81
37.71
ELF-REG-B (ours)
11.82
26.20
34.42
37.15
38.11
FMLM+ (Init) [ Agarwal et al., 2026 ]
5.10
15.10
26.10
31.80
33.60
DBTM [ Tang & Wang, 2026 ]
5.70
10.10
14.30
16.80
16.20
Appendix
Table 13: GSM8K mean pass@1 (%) across NFE. ELF-REG-L achieves the highest accuracy; at comparable parameter counts, ELF-REG-B also outperforms FMLM+ (Init) and DBTM at every NFE. FMLM+ (Init) and DBTM values are from their respective Table 8 [ Agarwal et al., 2026 ; Tang & Wang, 2026 ] ; ELF values are our evaluations. Bold marks the best value in each NFE column within each group.
Model
NFE / steps
MBPP-500
MBPP-378
HumanEval
HumanEval+
ELF-L baseline
4 NFE
4.20
8.10
5.68
5.30
ELF-L baseline
8 NFE
9.68
16.19
10.44
9.76
ELF-L baseline
16 NFE
12.89
20.42
14.10
13.38
ELF-L baseline
32 NFE
15.00
23.84
17.19
16.27
ELF-L baseline
64 NFE
15.95
25.69
18.33
17.49
ELF-L baseline
128 NFE
17.19
26.75
19.66
18.75
Appendix
Table 14: Code pass@1 (%) across inference budgets. ELF rows are our early-stop evaluations; external values are from PlaidQ Table 1 and Section 3.3 [ Peng et al., 2026b ] . Budget entries specify NFE or sampling steps; dashes denote unavailable measurements. Bold marks the maximum in each benchmark column over all methods and budgets.
Model
NFE / steps
MBPP-500
MBPP-378
HumanEval
HumanEval+
ELF-L baseline
4 NFE
16.30
26.84
16.39
15.25
ELF-L baseline
8 NFE
28.13
42.86
30.43
28.53
ELF-L baseline
16 NFE
34.75
47.57
35.67
34.15
ELF-L baseline
32 NFE
37.57
50.43
42.93
40.09
ELF-L baseline
64 NFE
38.41
52.88
45.04
42.28
ELF-L baseline
128 NFE
40.50
52.90
45.93
42.84
Appendix
Table 15: Code pass@10 (%) across inference budgets, using the same benchmark mapping and NFE / steps convention as Table 14 . The guided PlaidQ Table 1 row uses NFE 257. Bold marks the maximum in each benchmark column over all methods and budgets.
Model
NFE
pass@1
pass@2
pass@4
pass@8
pass@10
pass@16
MBPP-500
ELF-L baseline
128
17.19
24.10
31.28
38.34
40.50
44.80
ELF-REG-L (ours)
128
19.10
25.91
32.41
38.18
39.92
43.40
PlaidQ [ Peng et al., 2026b ]
129
12.19
–
–
–
21.44
–
PlaidQ+CFG [ Peng et al., 2026b ]
257
15.53
–
–
–
32.48
–
MBPP-378
Appendix
Table 16: Pass@ k (%) on MBPP-500, MBPP-378, HumanEval, and HumanEval+. ELF-L baseline and ELF-REG-L use NFE 128; we report sampling NFE for PlaidQ and PlaidQ+CFG. Bold marks the best result for each benchmark and sample budget across all methods.
Model
4
8
16
32
64
128
ELF-L baseline
5.47
7.58
8.56
9.65
9.65
10.55
ELF-REG-L (ours)
7.04
9.76
11.53
12.40
13.31
13.39
Appendix
Table 17: MATH-500 mean pass@1 (%) across NFE, ELF-L baseline and ELF-REG-L with early-stop generation. Bold marks the best value in each NFE column.
Model
pass@1
pass@2
pass@4
pass@8
pass@10
pass@16
ELF-L baseline
9.65
15.65
23.64
33.36
36.73
44.00
ELF-REG-L (ours)
13.31
20.75
29.95
40.41
44.04
52.20
Appendix
Table 18: MATH-500 pass@ k (%) at NFE 64, ELF-L baseline and ELF-REG-L with early-stop generation. Bold marks the best value at each sample budget.
Model
4
8
16
32
64
ELF-B baseline
28.67
44.00
45.23
45.32
45.34
ELF-REG-B (ours)
30.27
45.60
46.66
46.69
46.71
Appendix
Table 19: MMLU mean pass@1 (%), ELF-B baseline and ELF-REG-B with full-span generation. Bold marks the best value in each NFE column.
Setting
4
8
16
32
64
FMLM+ (Init) [ Agarwal et al., 2026 ]
5.10
15.10
26.10
31.80
33.60
FMLM+ (Init) (our reproduction) [ Agarwal et al., 2026 ]
Table 20: FMLM+ (Init) GSM8K mean pass@1 (%) across NFE. Reference values are from Table 8 in [ Agarwal et al., 2026 ] ; our two rows use 4 generation seeds.
ELF-REG-L (ours)
NFE
4
8
16
32
64
128
pass@10
31.48
46.39
49.99
52.78
54.98
55.05
Appendix
Table 21: Figure 1 (b): MBPP-378 pass@10 (%).
Hidden-state change ( ×10−4 )
Block
Condition
δedit
δ−
δ+
δ+−δ−
ν
t=0.05
4
baseline
1.562
0.120
0.123
0.003
0.016
4
REPA+REG shift-0
2.422
0.426
0.444
0.019
0.019
4
REPA+REG shift-1
1.867
0.407
0.449
0.042
0.045
12
baseline
0.176
0.068
0.095
0.027
0.117
Appendix
Table 24: Perturbation magnitude and direction on GSM8K at the supervised and final blocks. Values average paired single-latent interventions for the three ELF-B settings at epoch 6. REPA+REG shift-0 and shift-1 are ELF-REG-B (ours) settings. Shift-1 has a larger later-minus-earlier change than shift-0 at block 4 at each corruption time. At block 12, this ordering reverses at t=0.2 and 0.7 , while the difference at t=0.05 is small. δedit measures change at the edit, δ− and δ+ are means in matched earlier and later windows, and ν is the mean per-example normalized difference. Hidden-state change is 1− cosine similarity.
Condition
Validation pass@1 (%)
Test pass@1 (%)
ELF-B baseline
21.58
17.89
ELF-B baseline, Qwen3-1.7B-Base encoder
10.16
10.39
ELF-REG-B (ours) (default)
32.32
25.02
Text REPA during denoising: 0≤t≤0.4
31.05
25.17
REG target: mean pooling
27.05
23.65
REG only
25.39
17.97
Appendix
Table 25: GSM8K training ablations for ELF-B at epoch 6. Reference configurations appear in the top rows. Each remaining row follows the default REPA+REG setup (alignment depth 4, last-token REG target, response-only text REPA, and λREG=0.1 ), except for the stated difference in setting. The time interval restriction for REPA loss applies to text positions, excluding REG, on denoising examples; decoder-example alignment is unchanged. Values are single-seed pass@1 (%) at 65 NFE. Bold marks the best validation and test values across all rows.
Condition
Validation pass@1 (%)
Test pass@1 (%)
ELF-L baseline
33.50
27.37
ELF-REG-L (ours) (default)
44.04
34.04
Alignment depth: 6
42.09
33.13
Alignment depth: 10
41.21
33.28
Alignment depth: 11
40.04
33.13
Appendix
Table 26: GSM8K alignment-depth ablations for ELF-L at epoch 2.5. Reference configurations appear in the top rows. All settings follow the default ELF-L setup except for the stated difference in alignment depth; the default is block 8. Values are single-seed pass@1 (%) at 65 NFE. Other defaults and evaluation settings follow Appendix B.6 . Bold marks the best validation and test values across all rows.
ρ / NFE
4
8
16
32
64
1
7.25
15.75
25.27
30.95
34.19
2
10.60
23.22
30.89
34.08
35.84
4
11.62
25.75
33.43
35.69
37.62
8
11.82
26.20
34.42
37.15
38.11
16
11.22
25.82
35.30
37.55
38.32
Appendix
Table 27: GSM8K accuracy (%) for the early-stop ratio sweep, ELF-REG-B (ours) at epoch 12.
NFE
ELF-B baseline
ELF-REG-B (ours)
4
13.87
17.86
8
39.59
44.62
16
65.93
71.34
32
81.95
85.06
64
100.00
100.00
Appendix
Table 28: ELF-REG-B retains more shared solutions at every reduced full-span budget. Retention (%) is measured on the 4,693 question–seed pairs both ELF-B baseline and ELF-REG-B solve at NFE 64. Each lower budget uses its own full-span trajectory; retention at NFE 64 is 100% by construction.
A. Language generation
Method
Generation state
Response generation
Training or sampling approach
LangFlow [ Chen et al., 2026 ]
Token embeddings
Parallel sequence denoising
Flow matching with a learned noise schedule.
LDLM [ Meshchaninov et al., 2026 ]
Learned contextual latents
Parallel denoising, then token decoding
Joint training of the latent encoder, denoiser, and decoder.
ELF [ Hu et al., 2026 ]
Contextual encoder representations
Parallel denoising; final token decoding
Flow matching with a shared denoiser and decoder.
S-FLM [ Deschenaux & Gulcehre, 2026 ]
Hyperspherical token embeddings
Parallel sequence denoising
Riemannian flow matching on learned embeddings.
FMLM [ Lee et al., 2026 ]
Noisy one-hot token vectors
Parallel sequence generation
Distillation of a continuous flow into a flow map.
Appendix
Table 29: Selected methods for language generation and representation supervision. Panel A distinguishes the generated state, generation structure, and training or sampling approach. Panel B distinguishes intermediate-feature alignment from jointly denoised representations. AR denotes autoregressive generation.
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce the Continuous Embedding Diffusion Reasoner (CEDR), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We therefore learn compact representations from multiple layers of a strong autoregressive teacher. Their decomposition also enables asynchronous denoising at different rates. We show that prompt encodings need only preserve the information required for the correct text-conditional score, rather than exactly match teacher features, and use a staged curriculum to learn a compact prompt encoder that replaces the teacher Transformer at inference. We adapt DiffusionNFT to learned self-conditioning guidance and incorporate gold-solution endpoints to supplement sparse rewards. Our supervised models outperform reported results from recent continuous-diffusion baselines at comparable backbone scales on mathematical reasoning and HumanEval code generation. With a 638M-parameter denoising backbone and learned prompt conditioning, post-NFT CEDR-L achieves 63.74% pass@1 on GSM8K and 24.6% on MATH500 at 64 denoising steps, and 32.85% on HumanEval and 30.18% on HumanEval+ at 128 denoising steps. Code will be available at: https://github.com/chengxiang/CEDR.
Xiang Cheng
Duke University Department of Electrical and Computer Engineering
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today's leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs can be made effective with minimal adaptation to the discrete domain. We propose Embedded Language Flows (ELF), a class of diffusion models in continuous embedding space based on continuous-time Flow Matching. Unlike existing DLMs, ELF predominantly stays within the continuous embedding space until the final time step, where it maps to discrete tokens using a shared-weight network. This formulation makes it straightforward to adapt established techniques from image-domain diffusion models, e.g., classifier-free guidance (CFG). Experiments show that ELF substantially outperforms leading discrete and continuous DLMs, achieving better generation quality with fewer sampling steps. These results suggest that ELF offers a promising path toward effective continuous DLMs.
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continue reasoning after that text is cleared, using only a fixed-size carried state. We implement this state as a small number of register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks. We post-train dLLMs to decode a chunk of text, clear it while preserving the register values, and continue decoding from the prompt and carried state. In our main comparisons on LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains of up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation, where correct programs usually span several chunks. Finally, registers can be further refined with reinforcement learning on long-horizon reasoning tasks.
Albert Ge, Chandan Singh, Yufan Zhuang +3
University of Wisconsin–Madison · Microsoft Research · UC San Diego