Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length N; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy ε∈(0,1/8] and sufficiently large N, our sampler achieves seed-averaged total-variation error at most ε, with total masked-state submissions and sequential depth both bounded by O(NCε−a) for constants 0<C<1 and a>0. These guarantees use polynomial vocabulary size and an edge-response lower bound set by N and ε. The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires Ω(Ncεb) counterfactual submissions or commit rounds in the worst case, for constants c,b>0.
Figures & tables
Figure 1: Counterfactual edge probing; faint arcs show the hidden forest. Keeping the committed state unchanged, vary only source i between two hypothetical fillings and compare the predicted rows at masked readout j . With the same fully filled background, a response above the oracle-noise margin detects {i,j} . Section 4.2 packs and aggregates such tests.
Sampler / variant
Target / oracle
Submissions
Depth
Output
Sequential chain rule
Arbitrary / exact
N
N
Exact
One product batch
Arbitrary / exact
1
1
No error bound a
Anari et al. (2024)
Arbitrary / exact
O(N) , exp.
O(N2/3)
Exact
Anari et al. (2026)
Arbitrary / noisy b
O(NlogN) , exp.
O(N1/2)
TV
Li and Cai (2025)
TC/DTC / averaged prediction error
O(1+TC+DTC) (path: O(N) )
same
Mean KL
Chen et al. (2026)
Supplied T / exact
O(1+T) (path: O(N) )
same
Mean KL
Table 1: Sufficient bounds at fixed positive accuracy and polynomial vocabulary. Mean: seed-averaged divergence; exp.: expected resources (ours are pathwise). TC/DTC: total/dual total correlation; T : a supplied bound on either. Path: Proposition I.1 . Oracle conditions, accuracy dependence, and markers a,b : Appendices I – J .
Figure 2: Shared tests and majority recovery. (a) Each cell represents all bank/tail columns; each readout color has one chunk. Equal-color pairs (no diamond) get vote 0 without probing. Grey: nonseparation; red: illustrative errors. (b) Singleton peels contract the core per phase. (c) One centroid per component forms a terminal batch.
Figure 3: Total oracle submissions Qtot=Qpre+Qcf+R , including preprocessing, probes, and nonempty commit rounds. Points: six-run means; whiskers: min–max ranges (not standard errors or confidence intervals). Proposal settings are fixed per forest and N ; random budgets are selected per run on coupled curves. Proposal ranges ( ≤78 submissions) may lie inside markers. Dotted line: exact singleton Qtot=N , not a cap. Horizontal ticks: 1k=1024 .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Proof step
Key idea and conclusion
References
Construct a hard matching family
Nontrigger oracle replies hide dependence while each edge retains loss dedge≍η2 . The inclusion calculation checks the target assumptions and oracle error under the stated calibration.
Section C.2 ; Lemmas C.1 and C.2
Relate adaptive commits to output error
Later batches depend on sampled values. The midpoint identity expresses output affinity as an expectation of products of local affinities along these adaptive paths.
Section C.3 ; Lemma C.3
Limit what probes can reveal
The unresolved matching and triggers remain conditionally uniform. One candidate per source per submission gives EDgen≤NQ/m .
Section C.4 ; Lemmas C.4 and C.5
Force collisions with few rounds
Edge counting forces large active batches; a conditional Laplace bound then gives Z≳N/R with constant probability under the budgets in equation 46 .
Sections C.5 , C.6 and C.7 ; Lemmas C.6 , C.7 and C.8
Convert collisions into sampling error
Each collision contributes a factor 1−dedge . The midpoint identity gives a prior-averaged loss; averaging over seeds and selecting a worst-case instance yields the TV lower bound.
Sections C.8 and C.9 ; Lemmas C.9 and B.2
Recover the main-text lower bound
Calibrate the matching witness at the public parameters, then take the fixed-accuracy limit to obtain Theorem 3.2 .
Corollaries G.2 and G.3
Appendix
Table 2: Proof steps for the lower bound.
Figure 4: Three possible batch selections on an initially unresolved edge. Each panel shows the edge before (top) and after (bottom) one commit. Purple rings mark membership in Bt ; filled nodes are committed. Selecting both endpoints gives a collision; selecting one gives a crossing edge. Both events retire the whole edge, but in (b) j remains uncommitted. Selecting neither endpoint leaves the edge unresolved. The proof records the edge identity and both triggers in (a)–(b); the target matching itself is unchanged.
Figure 5: The coloring and three submitted states of Example D.2 , labeled by the steps of Algorithms 2 – 3 . All 15 positions are drawn separately at fixed locations. Fill denotes the color under τ ; square outlines mark masked readouts and double outlines mark sources. Small outer numbers are position indices, and inner symbols in (b)–(d) are submitted tokens. No vertex is committed by these probes. The hidden edges are shown only to explain the construction.
Proof task
Key argument and conclusion
Algorithm steps
References
Certify the vocabulary banks
Preprocessing gives a nonempty complement of Bi , whose tokens have true marginal mass below t , certifying the tail representative. Under the standard floor, (RF) bounds ℓi≤(4C/t)1/s .
Step 1
Lemma E.3
Recover low-degree neighborhoods
Separating colors reduce packed probes to one-source comparisons at a fixed boundary. RT–UEN and the screening margin make their votes correct; fresh-color majorities give Aj(HG)=NFG(j) for degree at most d , except on reached-screen failures of total probability at most δfail .
Step 2 – Step 3 ; Step 3.1 – Step 3.3
Lemmas E.2 , E.4 and E.5 ; Lemma E.6
Peel the high-degree core
Let nkhi count vertices of degree above d before phase k . On successful paths, the forest degree sum gives nk+1hi≤(4/d)nkhi . Either stopping test leaves maximum degree at most d , certifying the true residual forest.
Step 4 – Step 5
Lemma E.7
Complete safe parallel sampling
Singleton peels and one centroid per true residual component are safe. Exact-row products then equal joint batch conditionals. Centroid deletion halves component sizes, giving logarithmically many terminal rounds on successful paths.
Step 4 , Step 6
Lemmas E.2 , E.8 and E.9
Bound all-path resources
Count screen loops for Qcf , and peel and centroid commits for rounds. Guard caps rounds even on failed-screen paths. Each screen is one parallel stage: D≤1+R+Tscr . Qcf excludes preprocessing and commits.
All: Step 1 – Step 6 , plus Guard
Lemma E.8
Case (i): bound output TV
Each coordinate is committed once, so adaptive Hellinger composition bounds an analysis-only safe rule’s oracle error. Coupling to that rule adds at most δfail , giving seed-averaged TV at most εH+δfail .
All: Step 1 – Step 6 ; output-law analysis
Lemmas E.9 , E.10 and E.11 ; Theorem E.1
Appendix
Table 3: Proof steps for the upper bound. Step labels link to Algorithms 2 – 3 .
Figure 6: The successful-screen path of Example E.2 , with N=15 and d=9 . Orange and violet rings denote the next peel or centroid commit; pale crossed vertices are already committed. The fills here are neutral, not random colors. The step labels refer to Algorithm 2 . A fresh screen separates (a) from (b); the terminal forest then needs no further discovery. All vertices remain individually visible.
Figure 7: The separate contraction example of Example E.3 . The six orange-ringed hubs are committed one at a time in Step 4 , with no intervening screen or peel-set recomputation. Their 54 leaves are drawn individually, not bundled. The unpeeled center z changes from degree 10 to degree 4 . The lower panel keeps the same locations and crosses out only the committed hubs.
Family
Structure before relabeling
wF(N)
K⋆
Matching
Disjoint pairs
0.5
6.375289×10−9
Path
One path
0.25
2.801962×10−9
Binary tree
Complete binary tree
1/6
1.226162×10−9
Growing stars
At most Δ⋆(N) leaves per star
0.5/Δ⋆(N)
5.138930×10−12
Appendix
Table 4: Ideal target families and fixed absolute accuracy thresholds. Here Δ⋆(N)=max{16,⌈N⌉} , and K⋆=10−10L0 uses the reference in Appendix K.2 .
In this paper, we use random walks on graphs as a verifiable sandbox for studying parallel sampling strategies in masked diffusion models (MDMs). We train an MDM on random walk samples from a fixed graph. The graph and transition kernel are never shown to the model and serve as latent structure that is both controllable and enables evaluation. The framework provides a validity check for generated walks and a measure of distributional fidelity through the estimated transition kernel. Using simple graphs, we theoretically prove that parallel unmasking via widely used scores such as lowest entropy is not uniformly better than random parallel sampling; even with exact conditional probabilities, performance critically depends on the conditional dependence structure induced by the graph, a phenomenon difficult to isolate in benchmarks like Sudoku. We also develop training-free bisection samplers for MDMs, which take logarithmically many steps in the sequence length and are provably exact for random walks if the learned marginals are exact. Experiments on graph-walk tasks confirm that different parallel samplers perform better on different graph structures. Experiments on pretrained MDMs show that bisection-style samplers provide strong speed-quality tradeoffs on OpenWebText generation and reasoning benchmarks including GSM8K, MBPP, and HumanEval. Together, these results use graph walks to uncover conditional dependence as a key principle of parallel MDM sampling and translate this insight into efficient samplers that transfer to language generation and reasoning.
Inference-time scaling is a promising paradigm to improve generative models, especially when outputs must satisfy structural constraints or optimize downstream rewards. We consider Masked Diffusion Model (MDM) and introduce MDM-VGB, a discrete diffusion sampler that augments unmasking generation with theoretically principled reward-guided remasking. Inspired by the recent success of the classical Jerrum-Sinclair backtracking Markov chain in reward-tilted generation, MDM-VGB extends the backtracking random walk from a fixed prefix tree to a masked-state graph, allowing tokens to be unmasked and remasked at arbitrary positions. The resulting sampler favors unmasking and remasking moves that lead to higher-value partial configurations, enabling both effective high-reward generation and efficient repair of low-reward samples. We prove that MDM-VGB is robust to process-verifier noise and achieves quadratic complexity, while popular test-time heuristics such as best-of-N can incur exponential complexity due to error accumulation. Our theoretical findings are corroborated by strong empirical performance, particularly on popular constraint-satisfaction and scientific benchmarks such as Sudoku and QM9.
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
Zhenghao He, Bohan Liu, Guangzhi Xiong +1
Department of Computer Science, University of Virginia