Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models
Organizations: Lateral Intelligence
Abstract
We revisit a standard accepted practice in the continuous diffusion language model literature of fixing conditioning prompt tokens clean during training. We make a very simple modification: also noise the conditioning prompt tokens during training. We demonstrate that under this modified training objective, we achieve better generalization in combinatorial reasoning tasks such as Sudoku and N-Queens, with the largest gains on harder variants ( solve rate on Sudoku Hard), and increased diversity of generated solutions ( coverage on 10x10 N-Queens). We also show measurable improvements to natural language generation quality in modest dataset regimes with Gigaword summarization, but notably demonstrate that gains do not transfer to all natural language tasks (e.g open ended dialogue generation). Our method is a single line change to the training objective, requires no additional inference costs by default, and provides the flexibility of classifier-free guidance inspired guided sampling. Our \href{https://github.com/LateralIntelligence/noise-your-prompt} {code} is publicly available.
Figures & tables
| Easy ( clues) | Hard ( clues) | |||
|---|---|---|---|---|
| (uncond.) | ||||
| (length 64) | (length 100) | |||
|---|---|---|---|---|
| Accuracy | Coverage | Accuracy | Coverage | |
| (paste-in) | ||||
| Model | Dropout | Best Val NLL | Checkpoint Step | |
|---|---|---|---|---|
| M | (paste-in) | |||
| M | ||||
| M | ||||
| M | ||||
| M | ||||
| M | (paste-in) |
| Easy | Hard | ||||
|---|---|---|---|---|---|
| Setting | Mean | (improvement) | Mean | (improvement) | |
| — | — | ||||
| — | — | ||||
| Ablation | Setting | Accuracy | Coverage |
| Separate cond. time | Sudoku easy | — | |
| Sudoku hard | — | ||
| N-Queens | |||
| N-Queens | |||
| Per-token Bernoulli | Sudoku easy, | — | |
| Sudoku easy, | — |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Conditioning-time clean prob. | |||
|---|---|---|---|
| Difficulty | |||
| Easy | [64.95, 69.00] | [67.55, 71.60] | [66.60, 70.60] |
| Hard | [3.90, 5.75] | [4.55, 6.50] | [2.95, 4.60] |
| Sudoku | N-Queens | N-Queens | ||||
|---|---|---|---|---|---|---|
| Easy Acc. | Hard Acc. | Accuracy | Coverage | Accuracy | Coverage | |
| Train pairs | Best Val NLL | Checkpoint Step | |
|---|---|---|---|
| K | (paste-in) | ||
| K | |||
| K | |||
| M | (paste-in) | ||
| M | |||
| M |
| Accuracy | Coverage | ||||
|---|---|---|---|---|---|
| Setting | Mean | Mean | |||
| (length 64) | — | — | |||
| (length 100) | — | — | |||
| (length 144) | ||
|---|---|---|
| Accuracy | Coverage | |
| (paste-in) | ||
| Guidance weight | |||||||
|---|---|---|---|---|---|---|---|
| Setting | Space | ||||||
| Easy, | logit | ||||||
| prob | |||||||
| Hard, | logit | ||||||
| prob | |||||||
| Easy, | logit | ||||||
| (length 64) | (length 100) | |||
|---|---|---|---|---|
| Accuracy | Coverage | Accuracy | Coverage | |
| Group | Setting | Value |
|---|---|---|
| Model | hidden size / blocks / heads | / / (default), / / (Gigaword 132M) |
| cond. dim / dropout | / | |
| board length | (Sudoku), / (N-Queens / ), (Gigaword) | |
| Algorithm | pred. type | |
| loss type | flow | |
| (Sudoku), |