Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent's GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion's scores in both directions across tasks, while its AR parent's scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent's generation performance than in-place conversion.
Figures & tables
AR base
(i) full-FT
(ii) +align
(iii) block
(iii) × ent. §
(iv) uniform
HumanEval pass@10
22.97
1.01
1.99
1.51
0.00
0.00
GSM8K
12.59
1.67
2.12
1.82
0.15
1.59
MBPP
18.60
0.00
0.00
0.00
0.00
0.00
MMLU-Pro
19.80
15.20
15.20
15.40
0.60
8.40
BoolQ (%)
63.2
—
63.4
62.0
—
12.8
SQuAD-short EM (%)
67.4
—
44.0
22.0
—
0.6
Table 1: In-place conversion of OLMoE - 1B - 7B. Each arm uses 0.75–1 B supervised tokens and the fixed evaluation harness. The lower rows report the structured-output suite (BoolQ, SQuAD-short EM and CoNLL-JSON validity). We added this suite after the phase gates and ran it on the AR base and the three arms then live. The full fine-tuning arm had been superseded by (ii) at the first gate, and the entropy-decoder column reuses (iii)’s weights and was scored on the primary suite only. § The same block-model weights evaluated with the entropy-bound decoder; MMLU-Pro decreases from 15.4 to 0.6. Each cell uses one sampling run; pass@10 uses n=20 samples. Sampling seeds were not fixed.
Figure 1: The frozen-tower architecture ( Frozen-Tower-1B ). The frozen context tower is a bit-exact parent copy running causal attention over the committed prefix; the denoiser is a second copy of the same architecture, bidirectional over the masked canvas block, reading the tower through layer-aligned cross-attention ( Reda et al., 2026 ) , implemented as key/value concatenation into the denoiser’s own attention, so the attach introduces no parameters. One layer is expanded per tower; computation flows upward. Under the frozen-experts recipe only attention, router, and norms train (orange). Because experts and embeddings stay bit-identical to the parent, tower and denoiser alias them (green): the two-tower system holds parent +1.9 GB of weights at inference, under 2× parent (§ 6 ).
Figure 2: The output-length-dependent degradation, at a glance (all axes are percentages on a common 0 – 100 scale, frozen harness; ordered counterclockwise by output length). Both conversions stay close to the parent on the knowledge and short-output axes (BoolQ, MMLU-Pro, SQuAD), though InPlace-1B is ten points down on MMLU-Pro and SQuAD EM sits above the parent for the formatting reason given in § 3 . InPlace-1B (red) then collapses on the longer-output tasks (JSON validity 31.8, GSM8K 7.96, MBPP 0.4), while Frozen-Tower-1B (teal), the frozen-tower configuration, holds at or near the parent on the knowledge and structured-output axes, at 95% on GSM8K, and at 72–80% of it on the two code axes.
HumanEval pass@10
MMLU-Pro
frozen tower
64.04
54.2
trainable tower
63.83
40.8
difference
−0.21 †
−13.4 [ −19.5 , −7.3 ]
Table 2: Trainable versus frozen context tower. Same recipe, corpus, data order and training seed; frozen = the committed seed-1 checkpoint at 4,960 steps (about 490 M tokens), trainable at 5,059 steps (500.08 M). Differences are trainable minus frozen. † Committed-score difference; the paired re-decode comparison gives −0.33 [ −6.13 , +5.26 ] (Appendix A ). MMLU-Pro interval: independent-binomial approximation (Appendix A ).
Table 3: Cross-protocol evaluation. Two implementations on our task set, beside the published scores from the authors’ harness. Parent scores agree across all three implementations, whereas RND1 varies strongly with decoding mode. ∗ Parent results reproduced from Radical Numerics Inc. (2025) . ‡ Published MBPP and GSM8K results use sampled decoding on the authors’ harness; the parity cells use greedy decoding on our task set. For RND1, sampled pass@1 changes from 67.10 to 14.63 under native greedy decoding. The T=0.01 rows are read in the text; on HumanEval, pass@10 collapses to pass@1 there because n=20 near-deterministic draws buy no diversity. “n/r” marks a cell the source does not report.
attach density
HumanEval pass@10
GSM8K
MMLU-Pro
MBPP
25 % (12/48)
19.60
50.49
43.00
6.60
50 % (24/48)
40.09
66.72
57.00
35.40
75 % (36/48)
40.94
73.77
56.00
39.80
100 % (48/48)
47.11
77.33
54.60
43.60
Table 4: Interface ablations at a fixed tower. Four attach-density arms at 250 M supervised tokens with the same corpus, seed and harness. They differ in which tower layers the denoiser reads. The 25 % and 100 % arms were trained and scored on Hopper hosts, the 50 % and 75 % arms on a Blackwell host (Appendix K ); seed spread and unmeasured contrasts as in the text.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Context drift of the in-place arm after 1 B tokens. Left: per-layer relative Frobenius distance of InPlace-1B ’s keys and values from the parent’s on the same held-out tokens; the frozen arm sits at zero by construction. Right: cosine between the two residual streams.
InPlace-1B (in-place)
Frozen-Tower-1B (frozen tower)
mean loss, token-weighted
4.547
4.685
mean of per-sequence losses
4.619
4.758
paired difference, frozen − in-place
+0.139 (sd 0.581), in-place worse on 35.2% of sequences
parent initialisation, same objective (unregistered)
16.984
16.599
improvement from initialisation
12.44
11.91
Appendix
Table 5: Held-out objective of the pair. 25,000 held-out sequences, corruption seed 777, each arm under its own objective, in nats per scored token. The parent rows score the untrained initialisation of each arm under that arm’s objective. Two sequences in the frozen arm’s draw carried no scored token. They enter the token-weighted mean with zero weight, and the paired statistics use the remaining 24,998. The last two rows were added after the readings.
K
greedy pass@1
rounds per block
tokens per forward
commit entropy (nats)
1
53.05
32
1
0.33
2
0.00
16
2
1.53
4
0.00
8
4
1.65
8
0.00
4
8
1.66
16
0.00
2
16
1.92
32
0.00
1
32
2.42
Appendix
Table 6: The untrained witness against tokens committed per forward. Greedy pass@1 on HumanEval, masked keys suppressed, S=32 , 164 problems per cell. Sampled pass@10 ( n=20 ) was measured at K=1 (89.56) and K=4 (0.00). No band was placed on this curve.
Qwen3-4B-Base (parent)
Frozen-Tower-Dense
%
InPlace-Dense (in-place)
HumanEval pass@10
86.42
53.13
61.5
6.55
HumanEval greedy pass@1
59.76
37.80
63.3
0.61
GSM8K
81.88
69.90
85.4
5.38
MMLU-Pro
46.60
45.60
97.9
32.00
MBPP
67.60
41.00
60.7
0.00
BoolQ
86.80
86.60
99.8
80.20
Appendix
Table 7: Dense-parent matched comparison. Each arm uses 500 M supervised tokens with the same corpus, seed, and evaluation harness; each is evaluated with its native decoder. The “%” column gives the frozen-tower score as a percentage of the parent score. The prespecified replication criterion is met: rs=0.615≥0.45 and ri=0.076≤0.15 .
Figure 4: Retention against the parent’s generated length. Each point is one task (Table 8 ); the dotted line is the 256-token generation cap, at which generated length is censored. The trend is an association across tasks. Both in-place arms fall with generated length, from 0.69–0.99 at one token to at most 0.1 at 139 tokens and beyond; both frozen-tower arms stay near the parent until the longest tasks. SQuAD-short sits above 1 for the formatting reason given in § 3 .
30B MoE pair
4B dense pair
task
length
in-place
frozen
in-place
frozen
MMLU-Pro
1
0.82
0.99
0.69
0.98
BoolQ
1.1
0.99
1.00
0.92
1.00
SQuAD-short EM
7.7
1.22
1.61
0.53
0.94
CoNLL-JSON validity
18.3
0.32
0.98
0.10
0.98
HumanEval pass@10
138.6
0.07
0.80
0.08
0.61
Appendix
Table 8: The points of Figure 4 . Length is the parent’s mean generated tokens; 256+ marks the cap. Retention is arm over parent on the score the task is read on (JSON validity for CoNLL-JSON).
third by solution length
frozen tower ( Frozen-Tower-1B )
in-place ( InPlace-1B )
shortest third
81.65
13.43
middle third
70.28
0.00
longest third
50.18
0.31
Appendix
Table 9: HumanEval pass@10 by canonical-solution length , thirds of 55, 55 and 54 problems with median lengths of 1, 5 and 8 lines; each arm’s per-problem estimate is the mean of its three decode seeds.
cell
pass@10
95 % interval
frozen tower, 1 B ( Frozen-Tower-1B )
71.60
[65.6, 77.6]
in-place, 1 B ( InPlace-1B ), three-seed re-decode
4.61
[2.1, 7.7]
dense frozen tower, 500 M ( Frozen-Tower-Dense )
53.13
[46.5, 59.9]
dense in-place, 500 M ( InPlace-Dense )
6.55
[3.3, 10.2]
placement arm, 250 M
38.29
[31.8, 44.7]
1.7 B donor @ 1.136 tok/param
11.12
[6.9, 15.8]
Appendix
Table 10: Bootstrap 95 % intervals over the 164 HumanEval problems. Upper block: single cells with per-problem counts. Lower block: paired differences, resampling problems once for both arms; where a cell has three decode seeds, the per-problem estimate averages them. The in-place 1 B row uses the three re-decoded seeds of Table 13 , since the committed draw has no per-problem file.
arm
pass@10
HumanEval+ pass@10
pass@1 [95 %]
HumanEval+ pass@1
Frozen-Tower-1B (frozen tower)
71.60
63.29
33.54 [28.7, 38.5]
28.38
InPlace-1B (in-place)
6.19
6.19
1.65
1.49
Appendix
Table 11: HumanEval+ and pass@1 for the pair , from the committed generations. The in-place committed cell has no per-problem file, so its interval is not computed; its three decode seeds give pass@1 of 1.22–1.52 with intervals inside [0.3, 2.7].
cell
GSM8K [95 %]
MBPP [95 %]
frozen tower, 1 B ( Frozen-Tower-1B )
82.56 [80.5, 84.5]
53.40 [49.0, 57.8]
AR parent (Qwen3 - 30B - A3B - Base) ‡
88.02 [86.3, 89.8]
74.40 †
dense frozen tower, 500 M ( Frozen-Tower-Dense )
69.90 [67.5, 72.3]
41.00 [36.8, 45.2]
dense in-place, 500 M ( InPlace-Dense )
5.38 [4.2, 6.7]
0.00 [0.0, 0.0] §
placement arm, 250 M
64.29 [61.7, 66.9]
14.00 [11.0, 17.2]
1.7 B donor @ 1.136 tok/param
0.83 [0.4, 1.4]
0.60 [0.0, 1.4]
Appendix
Table 12: Bootstrap 95 % intervals over problems for GSM8K and MBPP , in percentage points, from the committed generations. † Re-scored one problem in 500 differently from the committed value; no interval. ‡ The parent’s committed figures differ between sources: the parity files carry 88.02 / 74.40, the results documents the main text quotes carry 87.19 / 74.60, and the re-scored generations give 88.02 / 74.60; every difference lies inside the interval shown. § No problem was solved in these cells, so resampling problems gives a degenerate interval. The exact binomial 95 % interval for 0 of 500 is [0.0, 0.7].
cell
metric
committed
draws (1234 / 2345 / 3456)
3-seed mean
range
frozen-tower 1 B, S=32 (Blackwell read)
HumanEval pass@10
71.60
68.36 / 66.60 / 67.46
67.47
1.76
frozen-tower 1 B, S=32 (Hopper read)
HumanEval pass@10
71.60
67.51 / 67.90 / 70.07
68.49
2.56
in-place 1 B
HumanEval pass@10
6.19
3.66 / 4.64 / 5.52
4.61
1.86
frozen-tower 500 M, seed 1
HumanEval pass@10
64.04
64.15 / 63.57 / 63.87
63.86
0.58
frozen-tower 500 M, seed 2
HumanEval pass@10
63.63
65.84 / 64.38 / 66.62
65.61
2.24
frozen-tower S=16
HumanEval pass@10
69.66
70.50 / 69.64 / 71.06
70.40
1.42
Appendix
Table 13: Decode dispersion. “Committed” is the single draw reported in the main text and is not revised. Greedy generations are token-identical across seeds in every cell (164/164 problems).
denoiser
HumanEval pass@10
MMLU-Pro
trainable params
tokens
tok/param
0.6 B donor @ 250 M
0.61
8.2
0.44 B
250 M
0.568
0.6 B donor @ 500 M
5.51
22.0
0.44 B
500 M
1.136
1.7 B donor @ 250 M
—
9.4
1.409 B
250 M
0.177
1.7 B donor @ 500 M
0.61
9.2
1.409 B
500 M
0.355
1.7 B donor @ 800 M
6.05
11.6
1.409 B
800 M
0.568
1.7 B donor @ 1.601 B
11.12
13.4
1.409 B
1.601 B
1.136
Appendix
Table 14: Denoiser size at a fixed tower and attach , ordered by tokens per trainable parameter (tok/param). The first four rows are the two donors’ own runs; the 800 M and 1.601 B rows are the extended run of the 1.7 B donor, whose own checkpoint at 0.355 tok/param (step 5,059) reads 12.0 on MMLU-Pro against 9.2 for the separate 500 M run, the two runs differing in learning-rate path (Appendix I ).
arm
seed
trained and scored on
kernel, block
log
final
gens
per-prob.
estimate
OLMoE (i) full fine-tune
42
GH200
uniform mask
–
n/v
–
–
SD
OLMoE (ii) frozen + align
42
GH200
uniform mask
–
n/v
–
–
SD
OLMoE (iii) block
42
GH200
compl., 32 → 64
–
n/v
–
–
SD
OLMoE (iv) uniform-state
42
GH200
uniform-state, 64
–
n/v
–
–
SD
InPlace-1B
42
B200
compl., 32 → 64
✓
rcpt
seeds
seeds
SD, +3s
Frozen-Tower-1B
42
GH200
blockwise, 32
–
rcpt
✓
✓
SD, +3s
Appendix
Table 15: Experiment inventory. The seed is the training data-order seed. Kernel: compl. = the independent per-position masking of the in-place arm, blockwise = the exactly- k masking of the two-tower kernel; the block entry is the canvas schedule. Marks: ✓ in the archived record; rcpt = off-host copy with a checksum receipt in the record; n/v = a final checkpoint that exists in the archive but whose off-host copy was not checksum-verified for this paper; – = not in the archived record; seeds = present only for the three decode-seed re-evaluations. Estimate: SD = committed single decode draw; +3s = three decode seeds also reported (Table 13 ). Every committed cell was scored on its training host.
arm
delivered budget (tokens / steps)
InPlace-1B
1,000,075,732 / 10,113
Frozen-Tower-1B
1.000 B / 10,113
Frozen-Tower-1B , training seed 2
499,236,606 / 5,050
S=16 stage
n/r
InPlace-Dense
499,313,300 / 5,050
Frozen-Tower-Dense
499,313,300 / 5,050
Appendix
Table 16: Delivered budgets. Supervised tokens and optimiser steps as recorded in each run-completion record, which identifies the evaluated final checkpoint. The frozen-tower 1 B run’s record states its budget as delivered exactly at 10,113 steps; it shares a bit-identical data manifest and step count with InPlace-1B , so the two arms trained on the same tokens. Nominal budgets are 250 M, 500 M and 1 B; n/r = not in the record; n/a = the arm did no training. Table 15 gives what the archived record holds for each arm. All arms were scored with eval_harness_v1 . The primary estimate for every reported number is the committed single draw, by the paper’s fixed reporting rule, and the three-seed mean is its uncertainty companion.