Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers. With identifier length and training budget fixed, we decode each model in several ways. Decoding alone moves a diffusion model's Hit@1 by 6.6 to 13.7 points. Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers. The generated identifier is right for 14-21% of NQ320K queries. We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes' probabilities. It matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead in Hit@1; on NQ320K, the lead comes from the model, not beam search. Starting from one sampled identifier, one-pass scoring removes 46-83% of masked diffusion's deficit to beam search; from generate-and-match, at most a quarter. On NQ320K, every paradigm largely memorises which identifier answers which query: random identifiers keep 83-90% of the Hit@1 of residual-quantised ones. There, product-quantised identifiers lead residual-quantised ones by 3.4 points in the autoregressive model and by -0.7 to +3.6 in diffusion models; across decodings, AR's gap exceeds diffusion's by 1.5-2.3 points, around our 2-point threshold. Paradigm comparisons must report each paradigm at its own recipe and best decoding.
Figures & tables
Table 1. AR leads or ties every learned-code column; on NQ320K, random codes keep most of each paradigm’s Hit@1. Hit@1 (%); AR: trie beam, diffusion: generate-and-match. rev./shuf.: RQ with reversed / shuffled levels. Bold: best model per column, not a significance claim; grey: not converged. For reference, on the same benchmarks, DDRO ( Mekonnen et al., 2025 ) reports (NQ320K / MS300K): BM25 14.1 / 18.9, SEAL ( Bevilacqua et al., 2022 ) 29.3 / 27.6, NCI ( Wang et al., 2022 ) 32.7 / 29.5, dense ANCE ( Xiong et al., 2021 ) 24.5 / 29.7, DDRO (PQ) 48.9 / 32.9.
Table 2. One-pass scoring matches or beats generate-and-match in 11 of 12 settings; chain-rule rescoring is mixed. Δ Hit@1 (points), first decoding minus second; grey: paired 95% interval (NQ320K; MS300K intervals not shown); bold: interval excludes 0.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
AR
MDLM, BD8, BD4
Initialisation
T5 1.1 base (C4)
DiT 12 layers, width 768, 12 heads, OpenWebText checkpoint (BD: one per block size)
cross-entropy on masked codes, unweighted; t∼U(0,1) per example (BD: per block)
Optimizer
AdamW, peak lr 3⋅10−3 , β2=0.95 , weight decay 0.01, gradient clip 1; linear decay, no warmup
AdamW, peak lr 1.5⋅10−2 (MDLM), 10−2 / 7⋅10−3 (BD, NQ320K / MS300K), β2=0.98 , weight decay 10−5 , gradient clip 1; 3% linear warmup, then constant
Budget
64K steps × 256
MDLM 32K × 512, BD 64K × 256; 16.4M examples in every cell
Appendix
Table 3. Settings of every grid cell. Learning rates come from a sweep on RQ codes and are reused for PQ and random codes. Every cell: mixed bfloat16, dropout 0.1, one RTX 6000 Ada; final weights evaluated (MDLM’s EMA unused).
NQ320K
MS300K
Documents
109,739
319,927
Training queries
307,373
367,013
Documents with a training query
108,026
319,927
Eval queries
7,830
808
Unseen documents / eval queries on them
1,713 / 1,755 (22.4%)
0 / 0
Training examples per cell
16M
16M
Appendix
Table 4. Corpora. Unseen: documents that are gold for no training query. Our MS300K is built with the DDRO script (title and content, truncated to 4,000 characters), ∼ 10% of MS300K Document; its eval set is DDRO’s dev set, not the standard 5,193-query dev set.
Table 6. Per-token accuracy (%) where AR and diffusion see the same context (NQ320K eval queries): level 1 from the query alone, level 16 with the other 15 levels given. Diffusion: range over MDLM, BD8, BD4. On training queries AR is near 100 at every level; diffusion PQ rises from 88–90 (level 1) to 90–97 (level 16), RQ falls from 88–89 to 79–85.
Table 7. NQ320K Hit@ k (%) split by whether the gold document is the gold of a training query (seen: 6,075 queries; unseen: 1,755 of 7,830). Each paradigm at its best decoding ( Table 2 ): AR trie beam, MDLM one-pass scoring, BD8 and BD4 chain-rule rescoring. Bold: best per column in the top block. g 10 generated queries per training document added at the same 16M examples, so real queries are seen ∼ 4.5 × less often; paired change against MDLM above, seen / unseen: −13.5 [ −14.8 , −12.1 ] / +1.1 [ +0.2 , +2.1 ].
Decoding
Generates
Docid chosen by
AR
greedy, no trie
1 code, free
exact match to a docid
greedy, trie
1 code, in the trie
the trie (always a docid)
beam 100, trie
100 codes, in the trie
summed log-probability
Diffusion
one sample
1 stochastic sample
code matching
Appendix
Table 8. The decodings compared, all on the same trained models. AR’s trie beam and diffusion’s generate-and-match give Table 1 .
Table 10. Error taxonomy (%). Outcome: share of queries by the rank of the gold document. Cause: share of rank-1 misses by what the wrong top-1 document is (first match wins, Section 4 ): near-duplicate of the gold, at most 2 of 16 levels wrong, topical, or unrelated. Shading grows with the share.
Id.
Size
Same title (%)
Empty text (%)
Cosine (random)
Topic
NQ320K
PQ
75
1
0
.745 (.041)
visa-policy pages
NQ320K
PQ
70
1
0
.226 (.041)
FIFA World Cup
NQ320K
PQ
59
2
0
.224 (.041)
shared Hinduism nav box
MS300K
RQ
512
92
100
.000 (.015)
empty pages
MS300K
PQ
737
92
100
.000 (.015)
empty pages
Appendix
Table 11. The largest groups of documents sharing one code. Id.: identifier; size: documents in the group; share with an identical title; share with an empty body; mean pairwise TF-IDF cosine (random pairs in parentheses); topic.
This paper shows how diffusion language models (DLMs) can be used as effective and efficient retrievers. Existing DLM-based retrievers (e.g., DiffEmbed) follow BERT-style encoding, representing each query or passage as a single mean-pooled vector. This ignores how DLMs are trained to generate responses through masked-position prediction under bidirectional attention, a capability that can provide stronger retrieval signals. We propose DiffRetriever, which uses the DLM's native masked-position prediction directly for retrieval. For each query or passage, DiffRetriever appends one or more masked positions, using the outputs as retrieval representations in a single forward pass. With one masked position, single-representation DiffRetriever already improves over DiffEmbed on the same backbones. DiffRetriever also naturally extends to multi-representation retrieval: DLMs process multiple masked positions jointly, enabling ColBERT-style fine-grained matching with little additional encoding latency. In autoregressive LLM retrievers, the same multi-representation strategy requires sequential decoding and therefore incurs much higher latency. DiffRetriever obtains the strongest aggregate effectiveness within our matched comparison, outperforming DiffEmbed, PromptReps, and RepLLaMA. Masked-position counts selected on training data transfer well across datasets, while per-query variation suggests headroom for adaptive allocation. Code is available at https://github.com/ielab/diffretriever.
Shuai Wang, Yu Yin, Shengyao Zhuang +2
The University of Queensland, Brisbane, QLD, Australia · CSIRO
Discrete diffusion language models generate text by iteratively denoising an entire response in parallel. At each step, they predict tentative tokens for every masked position, committing the confident predictions to the output and discarding the unconfident ones. We show that the discarded tokens are in fact a useful lookahead signal for retrieval-augmented generation: even low-confidence tokens often surface salient entities early in the denoising trajectory, enabling retrieval of stronger evidence before the output is finalized. We exploit this through Self-Augmenting Retrieval for Diffusion Language Models (SARDI), a dynamic RAG framework that uses these lookahead tokens to guide retrieval during denoising. SARDI is training-free, retriever-agnostic, and applicable to any reasoning-capable discrete diffusion language model. Across five multi-hop QA benchmarks, SARDI outperforms current training-free diffusion and autoregressive retrieval baselines at up to 8× higher throughput.
Paul Jünger, Justin Lovelace, Linxi Zhao +2
Department of Computer Science, Cornell University.
Generative retrieval (GR) ranks documents by autoregressively generating document identifiers. Because many GR methods rely on trie-constrained beam search, they are vulnerable to early pruning of relevant prefixes under finite-beam decoding. Planning Ahead in Generative Retrieval (PAG) mitigates this failure mode by using simultaneous decoding to compute a document-level look-ahead prior that guides subsequent sequential decoding. We reproduce PAG at inference time and stress-test its decoding behavior. Using the authors' released checkpoint and identifier/trie artifacts under the reported decoding setup, we reproduce the main effectiveness results on MS MARCO Dev and TREC-DL 2019/2020, and corroborate the reported beam-size-latency trade-off in our hardware setting. Beyond reproduction, we introduce plan drift diagnostics that quantify how intent-preserving query variations alter the planner's top-n candidate set and highest-weight planner tokens, and how these changes affect guided decoding. We find that PAG's planning signal is brittle under lexical surface-form variation: intent-preserving typos can trigger plan collapse, where the planned candidate pool shifts enough that the look-ahead bonus provides little useful guidance, effectively reverting decoding toward weaker unguided search. We further evaluate fixed-index cross-lingual robustness using non-English mMARCO queries against an English index, and assess query-side mitigation strategies that require no re-indexing; query translation provides the strongest recovery in our setting. Overall, our results confirm PAG's reported effectiveness and the benefit of planning-guided decoding under the released inference setup, while showing that these gains depend on the stability of the planning signal under realistic query variation and query-document mismatch.
Kidist Amde Mekonnen, Yongkang Li, Yubao Tang +2
University of Amsterdam Amsterdam, The Netherlands