Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
Organizations: University of Washington · Meta Superintelligence Labs
Abstract
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.
Figures & tables
| 2,9,13 Method | InfinityChat | Creative Writing | Image Gen. |
|---|---|---|---|
| (A) Randomize over ideas, not tokens | |||
| Gacha Decoding (ours) | |||
| Standard token sampling ( ) | 2.41 | 5.25 | 4.01 |
| Token sampling ( , ) | 3.28 | 3.07 | 4.66 |
| Token sampling ( , ) | 3.46 | 0.44 | 4.93 |
| Token sampling ( , ) | 3.19 | 0.31 | 4.90 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| \rectanglecolor gatedtint2-233-2 \rectanglecolor gatedtint2-533-5 \rectanglecolor gatedtint2-833-8 3,10,17,26 | InfinityChat | Creative Writing | Image Generation | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Quality- gated div. (raw div.) | Avg. quality of passing (raw qual.) | Gate pass rate (%) | Quality- gated div. (raw div.) | Avg. quality of passing (raw qual.) | Gate pass rate (%) | Quality- gated div. (raw div.) | Avg. quality of passing (raw qual.) | Gate pass rate (%) |
| Claude Opus 4.8 | |||||||||
| Gacha (ours) | (15.14) | 1.00 (0.80) | 79.5 | (11.92) | 17.4 (17.3) | 97.8 | (13.37) | 0.94 (0.82) | 57.6 |
| List | 7.87 (8.62) | 1.00 (0.89) | 88.7 | 5.65 (9.67) | 14.9 (11.8) | 19.0 | 6.78 (7.89) | 0.95 (0.85) | 66.2 |
| Verbalized | 4.10 (4.19) | 1.00 (0.95) | 95.4 | 6.91 (8.31) | 15.3 (14.1) | 46.8 | 3.58 (3.92) | 0.96 (0.87) | 67.6 |
| Persona | 5.54 (5.67) | 1.00 (0.94) | 93.8 | 8.37 (8.84) | 16.4 (15.7) | 83.8 | 4.63 (5.04) | 0.95 (0.87) | 74.9 |
| 3,9,15 | InfinityChat | Creative Writing | Image Generation | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | Greedy | Greedy | Greedy | ||||||
| GPT-OSS 120B | |||||||||
| Gacha Decoding (ours) | 4.6% | 2.5% | 4.5% | ||||||
| List | 7.62 | 4.71 | 38.1% | 1.93 | 0.06 | 96.8% | 7.84 | 4.48 | 42.9% |
| Verbalized | 6.07 | 3.83 | 36.9% | 4.49 | 0.65 | 85.5% | 4.73 | 3.19 | 32.7% |
| Persona | 4.64 | 4.64 | 0.2% | 6.71 | 6.58 | 1.8% | 4.74 | 4.67 | 1.5% |
| 2,9,13 Method | InfinityChat | Creative Writing | Image Gen. |
|---|---|---|---|
| (A) Randomize over ideas, not tokens | |||
| Gacha Decoding (ours) | |||
| Standard token sampling ( ) | 2.54 | 5.45 | 3.26 |
| Token sampling ( , ) | 3.07 | 2.56 | 4.25 |
| Token sampling ( , ) | 2.16 | 0.00 | 4.90 |
| Token sampling ( , ) | 1.79 | 0.03 | 5.02 |
| Model | Sampling | Reasoning | Max tok. |
|---|---|---|---|
| Qwen 3.5 family ( Qwen Team, 2026 ) | top- , top- , presence penalty 1.5 | on | 32,768 |
| GPT-OSS family ( Agarwal et al., 2025 ) | top- | medium | 65,536 |
| Inkling family ( Thinking Machines Lab, 2026 ) | top- | high | 65,536 |
| Gemma 4 family ( Team et al., 2026 ) | top- , top- | on | 32,768 |
| gpt-5.5 ( OpenAI, 2026 ) | API default | default | 32,768 |
| claude-opus-4-8 ( Anthropic, 2026 ) | API default | default | 32,768 |
| Target | Hotspot residues | Binder length |
|---|---|---|
| BHRF1 | A65, A74, A77, A82, A85, A93 | 80–120 |
| IFNAR2 | A45, A73, A75, A77, A89, A91 | 60–175 |
| PD1 | A33, A95, A97, A102 | 80–150 |
| TrkA | X294, X296, X333 | 50–120 |
| PDL1 | A37, A39, A49, A98 | 64–155 |
| DerF21 | A10 | 70–185 |