Pretrained autoregressive Transformers represent a large sunk investment in compute. Existing conversion methods reuse that investment by changing either the architecture (attention to recurrence) or the objective (next-token prediction to denoising), never both. We convert Qwen3 teachers at 1.7B and 8B into attention-free, bidirectional, gated-delta-rule diffusion students in three stages, so that each capability can be traced to the stage that kept or lost it. Language modeling transfers only partially and in-distribution; in-context retrieval does not transfer. On a multi-query recall probe where the teachers score 0.34-0.58, both converted students score 0.000, and diffusion pretraining alone does not restore retrieval. A retrieval curriculum in the final stage, which gradually lengthens the gap between a key-value table and the queries that address it, restores it only stochastically: on a fixed schedule, one seed in three learns to retrieve. Advancing the gap only while a running accuracy estimate stays above a threshold works for all three of those seeds, holds on real text, and carries unchanged to 8B, where two of three seeds succeed. The third had not learned within its fixed 16k-step budget: retrieval switches on abruptly at a seed-dependent step (6.5k and 11k in the other two), so a fixed budget can cut a late run off. One boundary survives every intervention: every model that learns retrieval scores 0.000 on tokens that never appeared in a retrieval episode, and an arm that resamples the key and value tokens every batch shows this is a coverage limit, not memorization of particular bindings. Separately, we convert a 7B code model into a 3:1 recurrent-attention block-diffusion hybrid over 85k steps and report two negative training results.
Figures & tables
1.7 B
8 B
teacher
Qwen3-1.7B
Qwen3-8B
student cells
gated delta rule
gated delta rule; RWKV-7 in one arm
stage-1 refinement
60k steps, batch 2 (30.7M tokens)
120k steps, batch 1 ( ≈ 30M tokens)
stage-1 agreement, in-sample
0.746
0.779
stage-1 agreement, held-out
not measured
0.544
stage-3 corpus
text8 / wikitext
text8
Table 1: Settings at both scales. Stage 3 uses batch 4, sequence length 256, learning rate 5×10−4 (the RWKV-7 arm: 1×10−4 , with gradient checkpointing) and a standard 16k-step budget at both scales (one 8B arm in § 4.4 used 24k). In-sample agreement is measured on the distillation text (§ 3.1 ). Cells with two values give the two 1.7 B arms (text8, wikitext); the 8 B run uses text8 with p=0.5 . p is the retrieval-episode probability. Wall times are on different hardware and are not comparable.
setting
lock-in
per-seed (trained)
disj.
1.7 B, open-loop
1/3
0.00 / 0.00 / 0.55
0.000
1.7 B, closed-loop
3/3
0.97 / 0.98 / 0.97
0.000
1.7 B, closed-loop, wikitext
1/1
0.955 – 0.995 (range)
0.000
8 B, closed-loop
2/3
0.90 / 0.78 /chance
0.000
Table 2: Conversion-scale replication. Retrieval on the trained token band, per seed (mean over the N∈{4,8,16}×gap∈{0,64,160} grid; marginal ≈0.023 ). The wikitext row is a single run; its cell gives the grid-cell range rather than a per-seed mean. disjoint is the identical probe with keys and values from a held-out token band. The open-loop row’s non-zero seed is a partial lock-in ( 0.40 – 0.70 across grid cells); the closed-loop schedule lifts the same three seeds to ceiling. Every row—including every solved one—reads 0.000 on the disjoint band.
Discrete diffusion language models generate text by iteratively denoising an entire response in parallel. At each step, they predict tentative tokens for every masked position, committing the confident predictions to the output and discarding the unconfident ones. We show that the discarded tokens are in fact a useful lookahead signal for retrieval-augmented generation: even low-confidence tokens often surface salient entities early in the denoising trajectory, enabling retrieval of stronger evidence before the output is finalized. We exploit this through Self-Augmenting Retrieval for Diffusion Language Models (SARDI), a dynamic RAG framework that uses these lookahead tokens to guide retrieval during denoising. SARDI is training-free, retriever-agnostic, and applicable to any reasoning-capable discrete diffusion language model. Across five multi-hop QA benchmarks, SARDI outperforms current training-free diffusion and autoregressive retrieval baselines at up to 8× higher throughput.
Paul Jünger, Justin Lovelace, Linxi Zhao +2
Department of Computer Science, Cornell University.
Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.
This paper shows how diffusion language models (DLMs) can be used as effective and efficient retrievers. Existing DLM-based retrievers (e.g., DiffEmbed) follow BERT-style encoding, representing each query or passage as a single mean-pooled vector. This ignores how DLMs are trained to generate responses through masked-position prediction under bidirectional attention, a capability that can provide stronger retrieval signals. We propose DiffRetriever, which uses the DLM's native masked-position prediction directly for retrieval. For each query or passage, DiffRetriever appends one or more masked positions, using the outputs as retrieval representations in a single forward pass. With one masked position, single-representation DiffRetriever already improves over DiffEmbed on the same backbones. DiffRetriever also naturally extends to multi-representation retrieval: DLMs process multiple masked positions jointly, enabling ColBERT-style fine-grained matching with little additional encoding latency. In autoregressive LLM retrievers, the same multi-representation strategy requires sequential decoding and therefore incurs much higher latency. DiffRetriever obtains the strongest aggregate effectiveness within our matched comparison, outperforming DiffEmbed, PromptReps, and RepLLaMA. Masked-position counts selected on training data transfer well across datasets, while per-query variation suggests headroom for adaptive allocation. Code is available at https://github.com/ielab/diffretriever.
Shuai Wang, Yu Yin, Shengyao Zhuang +2
The University of Queensland, Brisbane, QLD, Australia · CSIRO