Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.
Figures & tables
Figure 1: HC-DLM coupled diffusion process. The forward direction independently corrupts the continuous latent trajectory x0→xT and the discrete token trajectory k0→kT via q(xt∣xt−1) and q(kt∣kt−1) . The reverse direction couples the two channels: pϕ(xt−1∣xt,kt) advances the continuous state conditioned on the current token state, while pθ(kt∣xt) reads token distributions from the latent state, iterating until the clean states (x0,k0) are reached.
Figure 2: HC-DLM training pipeline. The encoder qψ(x0∣k0) maps clean tokens k0 to a continuous latent x0 , while independent forward kernels produce the noisy pair (xt,kt) . The token-conditioned denoiser dϕ(xt,kt,t/T) produces the clean-latent estimate x^0 (loss Jcont ), the token predictor pθ(k0∣x0) decodes x0 to k^0 (loss Jrecon ), and the encoder is regularized by Jent .
Method
#Params
Easy
Hard
Autoregressive
ARM (w/o ordering)
42M
9.73
–
ARM (with ordering)
87.18
32.57
Discrete diffusion
MDM (vanilla)
6M
6.88
3.62
MDM (top-prob.)
18.51
9.44
Table 1: Sudoku puzzle-solving accuracy (% ↑ ) on easy and hard splits. Best result per column in bold .
Method
#Params
CD4
CD5
Autoregressive
GPT-2 Scratch
6M
31.9
4.3
85M
45.8
5.1
303M
41.3
4.5
Stream-of-Search
250M
54.2
–
LLaMA
7B
41.1
6.7
Table 2: Countdown accuracy (% ↑ ). Best overall result per subtask in bold ; best result at the 6M parameter scale underlined .
Method
#Params
Gen. PPL ↓
Autoregressive
Transformer
108M
66.7
Diffusion
MDM
116M
103.9
SEDD
116M
115.9
Duo
116M
97.6
Table 3: Unconditional generation results on LM1B. Baseline results are from models retrained by the LangFlow authors ( Chen et al., 2026 ) .
Table 5: Ablation on Sudoku: discrete noise type and token-ordering heuristic. Accuracy (% ↑ ). Best result in bold .
Figure 3: (a) Accuracy of the decoded intermediate prediction x^0 at each normalized step t/T ; HC-DLM improves faster. (b) Final accuracy on Hard Sudoku vs. step count; HC-DLM stays robust with fewer steps.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Training time
LangFlow and retrained baselines †
∼ 292 h
Plaid
∼ 375 h
HC-DLM (ours)
∼ 48 h
Appendix
Table 6: Reported wall-clock training time for LM1B models.
NFE
Gen. PPL
Entropy
128
75.5
4.21
64
80.1
4.21
32
93.0
4.21
16
117.2
4.18
8
168.8
4.02
Ground truth
40.4
4.32
Appendix
Table 7: Sample quality across sampling budgets. Lower Gen. PPL is better.
Figure 4: Wall-clock LM1B sampling time per batch at 128 NFEs and sequence length 128. HC-DLM is faster from batch size 16 onward, with the gap widening as the batch grows. Lower is better.
Figure 5: Effect of token-channel classifier-free guidance. We report generative perplexity (lower is better) and entropy (higher is better) as the guidance scale w varies. The main LM1B result uses w=2.75 ; w=1 corresponds to sampling without guidance extrapolation.
Example
Generated text
Example 1
first 21 points of the quarter because they rallied by 21 points .<|endoftext|>It is , after all , Obama ’s favourite to win for president .<|endoftext|>Around 700,000 to 800,000 are required to pay their retails , including winter holidays , year-course , morning holidays , trips , and rigs taxes .<|endoftext|>Even with his lawyer , Dorothy Kinaitz , who has planned a separate appeal , he could soon appeal to the jury .<|endoftext|>Although Geoff Karlfeld ’s Irish blood was redbed by judges Alec Dominan and Dr. McGerson , the American men
Example 2
alcohol .<|endoftext|>This year , NDO has cut its profit forecast for the second quarter of 2009 .<|endoftext|>Ÿes , you are going to have a real team that is going to get the support and remain out of the race as to how very good they ’re going through , M̈cCain said .<|endoftext|>An aggressive Revolutionary Union militia , Tim Lappba , were killed in the fighting .<|endoftext|>The impact of the April quake was striking , with anticipated power managers warning from the city ’s summer thunderstorms .<|endoftext|>The Blue Millions Rural Care is likely to be offered if the county fails to comply .<|endoftext|>Publishing professionals
Example 3
.<|endoftext|>Ẅe have to work together to make sure that it is difficult to do all of the things that we don ’t do in the future , ḧe said .<|endoftext|>Meanwhile , a couple of elections in the northern port city of Yivazoum sent a complaint saying his support for Mr Tsvangirai and Mr Mugabe to a disastrous bid for a second term .<|endoftext|>Bill Greenberg , a Republican on the Finance Committee , said that if Congress was for delay , the State program linked to Bush ’s overhaul of the financial institutions was not upstate to slip out the majority .<|endoftext|>Undle Sarben
Appendix
Table 8: Unconditional samples generated by HC-DLM on LM1B.
M
Easy
Hard
1
81.73
55.81
2
91.59
67.34
4
93.75
71.11
8
71.82
45.52
16
78.80
51.89
Appendix
Table 9: Ablation on latent sequence length M (Sudoku, accuracy % ↑ ). All other hyperparameters follow the main configuration in Section B.4 ; M=1 is the setting used throughout the Sudoku and Countdown experiments, while LM1B defaults to M=128 , with one latent per token.. For a fair comparison of early-to-mid stage optimization efficiency, we report accuracy at the same checkpoint (560 epochs).
Figure 6: Wall-clock inference time versus denoising steps. Runtime per generated batch scales approximately linearly for both methods. Across tested step counts, HC-DLM is consistently faster, reflecting its efficient DiT implementation.
Method
Continuous variable
Update of the continuous variable
Input of the token predictor
Decoded token can change later
MDM
none
—
kt
no
VMD
global latent
not updated (sampled once)
kt,z
no
CADD
noised token embeddings
no learned update
kt,xt
no
CCDD
frozen token embeddings
learned denoiser
kt,xt
no
HC-DLM (ours)
global latent
learned denoiser
x^0 only
yes
Appendix
Table 10: Structural comparison with hybrid discrete–continuous diffusion models. MDM ( Sahoo et al., 2024 ) , VMD ( Zhang et al., 2025 ) , CADD ( Zheng et al., 2026 ) and CCDD ( Zhou et al., 2026 ) use the absorbing kernel, under which a token stays fixed once it is unmasked. For HC-DLM, x^0 is the denoiser’s clean-latent estimate.