Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives
Organizations: Edward · Data Science and Analytics Thrust, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China · Department of Computer Science, University of California, Los Angeles, USA · School of Computing, National University of Singapore, Singapore
Abstract
Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the number of unresolved positions. We introduce COFFEE, a plug-and-play framework that avoids this enumeration by separating sequence dependence from the objective. At each diffusion step, a target-free carrier absorbs the marginal token distributions predicted by the denoiser to construct a joint model over the unresolved tokens, while a compiled finite-state model records how their combinations affect the sequence-level preference. Pairing their states allows COFFEE to transfer global preferences to unresolved positions and sample a clean reconstruction without retraining the diffusion model. The same framework supports explicit hard constraints and learned soft objectives. We evaluate COFFEE across multiple symbolic, language, and biological benchmarks, where it achieves strong control results with task-dependent quality and diversity trade-offs. By making objectives available to inference rather than only evaluation, COFFEE brings joint conditioning, completion-weighted guidance, and optimization-based constraints into pretrained neural generation, showing the potential of neural-symbolic methods in diffusion guidance.
Figures & tables
| Dyck repair | Sudoku | CommonGen | RTP | Dyck repair | Sudoku | CommonGen | RTP | DeepSTARR | APARENT | K562 | Protein | |||||||||||||
| Method | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality | Control | Quality |
| Base | 3.0 | 13.5 | 60.4 | 0.1 | 38.6 | 10.1 | 53.5 | 5.2 | 7.5 | 14.0 | 20.4 | 0.3 | 40.9 | 14.4 | 43.3 | 5.3 | -0.4 | 728.3 | 2.3 | 728.3 | 44.3 | 728.3 | 11.0 | 1.7 |
| CDD | 46.9 | 14.8 | 99.3 | 0.0 | 46.1 | 30.4 | 55.7 | 17.2 | 62.5 | 14.7 | 90.7 | 0.0 | 40.9 | 18.8 | 81.7 | 13.1 | 3.5 | 1.2 | 7.9 | 682.1 | 83.7 | 70.1 | 10.5 | 1.7 |
| CDM | 0.2 | 11.7 | 100.0 | 0.0 | 57.2 | 33.8 | 65.6 | 5.4 | 7.7 | 13.8 | 100.0 | 0.0 | 41.5 | 14.5 | 54.3 | 4.9 | 0.6 | 429.1 | 10.1 | 691.9 | 72.4 | 250.9 | 10.2 | 1.7 |
| D-CBG | 14.9 | 14.3 | 99.8 | 0.0 | 39.4 | 10.4 | 62.2 | 5.0 | 23.1 | 14.5 | 98.7 | 0.0 | 30.7 | 32.7 | 72.4 | 7.1 | -0.5 | 1799.7 | 2.2 | 1492.4 | 100.0 | 1.2 | 11.0 | 1.7 |
| DG-TAG | 100.0 | 14.9 | 100.0 | 0.0 | 29.7 | 9.8 | 87.9 | 7.2 | 100.0 | 14.9 | 100.0 | 0.0 | 32.9 | 14.7 | 66.9 | 5.4 | 3.6 | 2.6 | 24.3 | 6.3 | 97.2 | 3.4 | 14.5 | 1.5 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Instance | Product state | Reward factor |
|---|---|---|
| Diffusion-only | identity | |
| Carrier-only | ||
| Learned guidance | ||
| Hard constraint | ||
| Strict repair |
| Task / backbone | Base | CDD | CDM | D-CBG | DG-TAG | GILC-DB | MDM-VGB | Coffee |
|---|---|---|---|---|---|---|---|---|
| Dyck / Dream | 6.249 | 96.966 | 51.076 | 82.366 | 121.766 | 443.709 | 406.640 | 13.940 |
| Dyck / LLaDA | 6.237 | 107.607 | 76.587 | 107.587 | 136.590 | 469.933 | 389.544 | 12.086 |
| Sudoku / Dream | 50.592 | 105.776 | 67.333 | 95.105 | 83.739 | 111.114 | 53.027 | 56.295 |
| Sudoku / LLaDA | 69.252 | 110.038 | 81.719 | 110.513 | 101.191 | 112.844 | 71.603–72.364 | 75.136 |
| CommonGen / Dream | 20.121 | 38.690 | 111.330 | 20.272 ‡ | 80.754 | 30.454–30.649 | 624.670 | 109.432 |
| CommonGen / LLaDA | 15.428 | 68.535 | 152.613 | 79.558 | 76.392 | 62.749 | 590.824 | 65.442 |
| Component | Formal statement | Empirical evidence |
|---|---|---|
| Backward DP | finite exact recursion | graph/support sizes |
| Ancestral replay | exact at | sampling rule and seed |
| Scorer fit | convex fixed-design objective | solver status and validation metrics |
| Carrier fit | no global-optimum claim | epoch/loss/convergence flag |
| Diffusion host | frozen | checkpoint and schedule |
| Task | Scored scope | Automatic controller construction | Terminal rule |
|---|---|---|---|
| RTP | generated continuation only | Mine and select train-only count features for ; compile signed token terms and selected phrase weights into one weighted AC machine. | zero for every state |
| Common Gen | generated sentence | For each record, enumerate standalone base, regular-plural, and possessive forms; tokenize them and compile the reachable lexical-prefix state plus a satisfied-concept bit mask. | iff every required concept is complete and no forbidden concept is hit; otherwise |
| Dyck repair | 12 locked prefix tokens plus 20 repaired suffix tokens | Compile a bounded-stack graph, intersect forward/backward min-plus optima, then attach frozen evidence and carrier weights. | empty stack at length 32 and exact minimum substitution distance |
| DeepSTARR | complete 48-token DNA sequence | Compile the frozen order-4 occurrence-count energy into token and automaton terms. | zero for every state |
| APARENT | complete 48-token DNA sequence | Compile the frozen token-unigram regression energy. | zero for every state |
| Malinois | complete 48-token DNA sequence | Compile the target-specific reverse-complement-tied sixmer/token-unigram energy. | zero for every state |
| Task | Generation problem | Target / constraint | Primary control metric | Quality and audit metrics |
|---|---|---|---|---|
| RTP | prompt continuation | non-toxic continuation | ToxicBERT continuation toxicity and satisfaction | prompt-conditioned GPT-2-large corpus PPL |
| Common Gen | concept realization | terminal lexical DFA acceptance | fail-closed all-concept satisfaction | exact/normalized coverage; forbidden hits; unreachable support; conditional PPL |
| Dyck | fixed-length repair | valid minimum-edit completion | exact validity | all-output Hamming edit and legal-edge rate |
| Sudoku | clue-preserving completion | valid grid | exact validity | clue preservation and output format |
| DeepSTARR | unconditional DNA generation | developmental activity | official DeepSTARR H5 predictor | native masked LOO and failure count |
| APARENT | unconditional DNA generation | isoform usage | official APARENT H5 predictor | native masked LOO and diversity diagnostics |
| Method | Implementation | Fitted object and supervision | Online rule | Reported compute |
|---|---|---|---|---|
| Base | host implementation | none | native reverse diffusion | reverse updates |
| CDD | same-backbone port | train-only projection expert | top- projection with augmented-Lagrangian inner updates | plus inner-loop count |
| CDM | sequential Monte Carlo (SMC) port | train-only twist | particles with fixed proposal, effective sample size (ESS) threshold, and resampling law | task-dependent particle work; see below |
| D-CBG | same-backbone port | train-only corrupted-state classifier | classifier reweights token alternatives; DNA uses the declared full-vocabulary Taylor form | updates and candidate width |
| DG-TAG | method-inspired port | frozen task predictor | rescore top- successor states and update logits | updates and predictor calls |
| GILC-DB | adapted port | train-only task reward model | Monte-Carlo logit correction at every reverse update | reverse updates; MC samples reported separately |
| Base | Coffee | |
|---|---|---|
| 6.976 | 6.809 | |
| 27.903 | 27.238 |
| Field | COFFEE language default |
|---|---|
| Diffusion checkpoint | LLaDA-8B-Base |
| Generation / block length | 32 / 32 |
| Denoising budget | 32 |
| Evidence temperature | 1.0 |
| Carrier-emission temperature | 1.0 |
| Ancestral temperature | 1.0 |
| Benchmark | Development data | Final panel | Frozen selection rule and interpretation |
|---|---|---|---|
| CommonGen | 993 records, with 992 in the corrected common pool | 1,497 disjoint records | Maximize exact lexical acceptance under the common PPL cap, followed by coverage, PPL, and strength tie breaks. |
| RTP | Method-independent development panel | Common 1,024-prompt stress panel selected by Base toxicity | Maximize non-toxicity under the same-budget quality rule. Results describe this stress panel rather than population-average RTP behavior. |
| Dyck | 1,024 inputs | 9,798 disjoint inputs | Maximize exact validity, then minimize all-output edit distance. Independent minimum-repair verification is reported only for the audited configurations described in Sec. H.2 . |
| Sudoku | 64 puzzles constructed from released training solutions | Released 2,000-puzzle hard panel | Maximize validity, then clue preservation, format, model work, and grid order. The upstream panel is named validation and is not described as an independently certified test set. |
| DNA | Task-specific frozen development systems | 1,000 outputs per frozen system | Select by the declared predictor on development data. Native LOO and diversity remain separate diagnostics. Comparisons are distributional rather than molecule paired. |
| Protein-1YCR | Common development geometries | Common 400-geometry panel | Maximize Success@1Å, then minimize mean motif RMSD and evaluator work. Particle methods retain their additional model and evaluator calls. |
| Method | Parameters searched on development data | Computation fixed before final evaluation |
|---|---|---|
| COFFEE | Language evidence, carrier, and ancestral temperatures in , followed by reward strengths in for declared representatives. Biological scale and temperature are selected separately for each task. | Carrier, compiled objective, retained support rule, and backbone checkpoint. |
| CDD | Projection strength , initial multiplier , and inner iterations where applicable. | Projection expert, candidate width, and outer update schedule. |
| CDM | Twist strength for language, with task-specific biological refinements. | Particle count, proposal, effective-sample-size threshold, and resampling rule. |
| D-CBG | Classifier strength , with separately selected full-vocabulary scales for Malinois. | Corrupted-state classifier and candidate rule. |
| DG-TAG | Rescoring strength , plus the declared CommonGen refinement. | Frozen task predictor and top- successor set. |
| GILC-DB | Language strength , with task-specific biological scale refinements. | Reward model, Monte Carlo sample count, and candidate width. |
| Main-text result | Additional evidence | Required interpretation |
|---|---|---|
| Dyck repair | The compiler constructs the minimum-substitution valid support before probabilistic replay. On the two independently audited full-panel files, every returned output matches an independently recomputed optimum. | The independent audit covers the stated Dream and LLaDA configurations. It is not a blanket empirical certificate for every budget. |
| Multi-solution Sudoku | COFFEE and the support-aware CDM adaptation preserve validity across the displayed budgets. Sec. H.1 separately compares coherent joint replay with independent exact marginals. | Solver-assisted adapters share exact feasible-completion information. Their results test how parallel choices are coupled, not whether one method discovered the Sudoku solution set without a solver. |
| CommonGen and RTP | CommonGen uses a record-conditioned hard automaton. RTP uses a learned soft objective on a common 1,024-prompt stress panel. | Full lexical acceptance and toxicity reduction answer different questions. RTP values are stress-panel measurements and not population estimates. |
| Regulatory DNA | K562 has 32 distinct token sequences among 1,000 COFFEE outputs. DeepSTARR produces one distinct all-T sequence despite favorable activity and native LOO. | Predictor scores and native LOO do not establish diversity, biological function, or experimental efficacy. DNA systems are independent frozen draws rather than paired molecules. |
| Protein-1YCR | Success@1Å and mean motif RMSD are computed on the common 400-geometry panel. | Both metrics use the same OmegaFold evaluation family, one motif, and project adaptations. They are not independent wet-lab validation. |
| Method | Full ( ) | ||||
|---|---|---|---|---|---|
| Joint replay | 120 | 120 | 120 | 120 | 120 |
| Independent exact marginals | 68 | 100 | 114 | 118 | 120 |
| Uniform joint | 120 | 120 | 120 | 120 | 120 |
| D-CBG | 13 | 40 | 71 | 113 | 120 |
| Joint reference | 120 | 120 | 120 | 120 | 120 |
| Maximum cells committed per step | 14 | 7 | 4 | 2 | 1 |
| Carrier and sampling rule | Lexical coverage (%) | PPL |
|---|---|---|
| H–Neutral | 22.01 | 16.81 |
| H–Feasible | 100.00 | 20.28 |
| H–Full | 100.00 | 18.93 |
| H–Full-marginal | 83.98 | 18.34 |
| I–Neutral | 30.99 | 19.56 |
| I–Feasible | 100.00 | 23.71 |