We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.
Figures & tables
Figure 1: We introduce Gacha Decoding , an inference-time method that strongly improves LM response diversity ( A ). We recast diversity as an instruction-following problem: instead of relying on token entropy, Gacha pairs a model’s IF capabilities with external RNG to identify and realize distinct output modes. This reformulation enables Gacha to improve with stronger models ( B ), even when diversity under prior approaches stagnates or declines. ( A–B ): results on InfinityChat ( Jiang et al., 2026 ) . ( C ): random sampled images from text prompts generated by ours and baselines.
Figure 2: Gacha Decoding generates responses by “planning with dice.” Given a task, we instruct the LM to identify and make a sequence of decisions to determine the final response. At each step, the LM enumerates possible options, and uses external randomness to select one uniformly. This design utilizes the LM’s IF capabilities while minimizing reliance on token entropy.
Figure 3: Gacha Decoding achieves much higher quality-controlled diversity than prior approaches across all models and task domains, with gains increasing at larger generation scales. This gain in Vendi translates to strong improvements in mode-coverage sample efficiency (dashed lines). For example, on open-ended chat with Opus 4.8, Gacha requires 11.0x fewer samples to cover the same number of modes as 256 samples from persona prompting (next-best prior approach).
Figure 4: Our method derives diversity from an LM’s instruction following ability as opposed to token entropy. Hence, our method is the only approach to consistently improve as the underlying LM’s capabilities do, even when token entropy and diversity under entropy-reliant methods decline.
Figure 5: We use GPT 5.5 with Gacha Decoding to design diverse protein binders against downstream targets. Gacha yields the most structurally-diverse binders (TM-score Vendi) among designs passing a binding-quality threshold (ipSAE min>0.5 ), with the gap widening with design budget. Against PD-1, Gacha Decoding discovers binders with mixed α / β secondary structure ( β -sheets colored cyan), while Verbalized Sampling remains limited to all-helical folds.
2,9,13 Method
InfinityChat
Creative Writing
Image Gen.
(A) Randomize over ideas, not tokens
Gacha Decoding (ours)
10.08
9.44
9.23
Standard token sampling ( T=1 )
2.41
5.25
4.01
Token sampling ( T=3 , K=10 )
3.28
3.07
4.66
Token sampling ( T=3 , K=20 )
3.46
0.44
4.93
Token sampling ( T=5 , K=10 )
3.19
0.31
4.90
Table 1: We compare Gacha Decoding against ablations of our three design choices ( section 2 ), measuring diversity as quality-gated Vendi score ( ↑ , n =128). We ablate (A) idea-space sampling via extreme-temperature top- K token sampling, (B) external randomization over LM-enumerated candidates via no-RNG and no-enumeration variants, and (C) sampling ideas autoregressively as a series of conditional local decisions by instead generating complete ideas directly (idea-list) or sampling decisions independently (independent axes). All ablations drop diversity—often sharply. Results on (A) are from GPT-OSS-120B, where temperature is configurable; (B, C) show Opus 4.8.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Gacha Decoding can discover novel modes, surfacing several stories narrated in the second person when no other approach does so across 256 samples. Pronouns are bolded for emphasis. The task prompt, “Write me a short horror story about singapore apartment house,” is drawn from WildChat ( Zhao et al., 2024 ) and included amongst our story evaluation prompts.
Table 2: Extended main results for our experiments in section 3 , showing the concrete quality and diversity scores underlying fig. 3 . Quality-gated diversity : Vendi score over the samples that pass a quality gate, our main metric. We show raw Vendi over all samples—including low-quality ones—in parentheses. Average quality of passing samples : mean judge quality of the passing samples (chat response validity, binary 0/1; story rubric-judged quality, 0–20; image rubric-judged alignment, 0–1), with the mean over all samples in parentheses. Quality-gate pass rate : the percentage of samples that pass the quality gate, defined to be 0.95× the model’s direct-sampling mean on the same prompt. We show results at the highest number of samples per prompt considered for each model ( n=512 for open-weight models GPT-OSS and Inkling, and n=256 on closed models).
3,9,15
InfinityChat
Creative Writing
Image Generation
Method
T=1
Greedy
Δ
T=1
Greedy
Δ
T=1
Greedy
Δ
GPT-OSS 120B
Gacha Decoding (ours)
10.08
9.62
− 4.6%
9.44
9.20
− 2.5%
9.23
8.82
− 4.5%
List
7.62
4.71
− 38.1%
1.93
0.06
− 96.8%
7.84
4.48
− 42.9%
Verbalized
6.07
3.83
− 36.9%
4.49
0.65
− 85.5%
4.73
3.19
− 32.7%
Persona
4.64
4.64
+ 0.2%
6.71
6.58
− 1.8%
4.74
4.67
− 1.5%
Appendix
Table 3: We show quality-gated Vendi score diversity ( ↑ higher is better, n=128 ) with Gacha Decoding and other inference-time methods under (a) standard token sampling ( T=1 ) and (b) greedy decoding. Because Gacha Decoding externalizes randomness from the LM and does not rely on token entropy, switching to greedy decoding only negligibly affects the diversity it elicits ( Δ ). The near-complete drop in Verbalized Sampling and list prompting diversity on creative writing occurs because these methods often degenerate into nonsensical responses under greedy decoding, such that few if any responses pass the quality gate.
2,9,13 Method
InfinityChat
Creative Writing
Image Gen.
(A) Randomize over ideas, not tokens
Gacha Decoding (ours)
11.65
8.44
9.04
Standard token sampling ( T=1 )
2.54
5.45
3.26
Token sampling ( T=3 , K=10 )
3.07
2.56
4.25
Token sampling ( T=3 , K=20 )
2.16
0.00
4.90
Token sampling ( T=5 , K=10 )
1.79
0.03
5.02
Appendix
Table 4: Additional ablation results to corroborate table 1 . Panel (A) shows results with Inkling Large; panels (B) and (C) show GPT-5.5.
Model
Sampling
Reasoning
Max tok.
Qwen 3.5 family ( Qwen Team, 2026 )
top- p=0.95 , top- k=20 , presence penalty 1.5
on
32,768
GPT-OSS family ( Agarwal et al., 2025 )
top- p=1.0
medium
65,536
Inkling family ( Thinking Machines Lab, 2026 )
top- p=1.0
high
65,536
Gemma 4 family ( Team et al., 2026 )
top- p=0.95 , top- k=64
on
32,768
gpt-5.5 ( OpenAI, 2026 )
API default
default
32,768
claude-opus-4-8 ( Anthropic, 2026 )
API default
default
32,768
Appendix
Table 5: Models used in our experiments with their sampling parameters. Unless otherwise noted for specific experiments, all models are sampled at T=1.0 . For open-weight models we adopt each model card’s recommended generation settings with reasoning enabled; for API models we use the provider defaults (Opus 4.8’s default reasoning is adaptive, high).
Target
Hotspot residues
Binder length
BHRF1
A65, A74, A77, A82, A85, A93
80–120
IFNAR2
A45, A73, A75, A77, A89, A91
60–175
PD1
A33, A95, A97, A102
80–150
TrkA
X294, X296, X333
50–120
PDL1
A37, A39, A49, A98
64–155
DerF21
A10
70–185
Appendix
Table 6: Targets considered in our protein binder design pilot ( Pacesa et al., 2025 ) . Hotspots are specified by chain ID and residue number in the input structure. Binder lengths are sampled uniformly from the ranges shown (inclusive).
School of Information Science and Technology, ShanghaiTech University, Shanghai, China · State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China.