Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation
Authors: Alireza Jafari, Arman Adibi, Mohammad Ghavamzadeh, Hadi Daneshmand
Organizations: Department of Computer Science University of Virginia Charlottesville, VA, USA · School of Computer and Cyber Sciences Augusta University Augusta, GA, USA · Qualcomm AI Research San Diego, CA, USA
Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an ε-Nash equilibrium in O(1/ε) time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to 18× larger, without any fine-tuning or retraining, at the cost of additional test-time computation.
Figures & tables
Figure 1: Revision Game. Each token position is a player whose choice is evaluated given all context tokens. Autoregressive decoding leaves earlier tokens fixed, whereas the tokens are revised in the revision game. We used GPT-2 small for conditional probability estimates. At a Nash equilibrium, no player can increase its conditional probability by changing its token. Details are reported in Appendix A .
Figure 2: Convergence of Nash decoding on WikiText-103. (a) Running minimum of the Nash gap, compared with a 1/k reference curve. The reference indicates a rate of decay, aligned with Theorem 1 . Both axes are log-transformed. (b) GPT-2 XL continuation perplexity and ModernBERT pseudo-perplexity. The x-axis is log-transformed.
CoQA
PubMedQA
CLAPNQ
Model
Size
F1
R-L
R-Ls
F1
R-L
R-Ls
F1
R-L
R-Ls
One-shot masked LMs (not fine-tuned)
RoBERTa-L ( Liu et al., 2019 )
355M
28.00
31.45
31.46
7.99
8.21
8.74
5.94
10.38
10.59
ModernBERT-L ( Warner et al., 2025 )
395M
39.54
42.92
42.93
5.93
7.95
8.34
7.03
12.05
12.66
Ettin-400m ( Weller et al., 2025 )
396M
43.87
46.22
46.24
4.03
4.55
4.76
6.97
11.78
12.41
Autoregressive models
Table 1: Results across three question-answering benchmarks under a shared evaluation protocol. Size denotes the number of model parameters; for Nash decoding, it denotes the masked model size. Best results are bolded and second-best results are underlined . The complete benchmark results are reported in Appendix F .
Figure 3: F1 score and Nash gap during Nash decoding iterations on CLAPNQ.
Model
Size
F1
R-L
R-Ls
Left-to-right initialization
mmBERT-base
308M
4.22
6.11
5.38
RoBERTa-Large
355M
22.23
19.79
20.77
ModernBERT-Large
395M
21.91
19.88
19.99
Ettin-400m
396M
17.41
17.29
17.36
All-mask initialization
Table 2: Nash Equilibria Basins on PubMedQA.
CoQA
PubMedQA
CLAPNQ
Decoding rule
F1
R-L
R-Ls
F1
R-L
R-Ls
F1
R-L
R-Ls
Uniform-ordered
38.65
40.05
40.08
30.50
23.51
27.14
56.15
50.08
55.00
Confidence-ordered
41.85
42.21
42.28
31.74
24.64
28.50
58.84
54.30
58.11
Max Nash gap
43.07
43.22
43.31
32.11
24.80
28.75
59.16
54.67
58.37
Table 3: Nash decoding versus single-token masked-diffusion decoding using the same masked model. All rows use the instruction-tuned ModernBERT-chat with its corresponding prompt; details in Appendix H .
Figure 4: Runtime comparison of Nash and autoregressive decoding on CoQA. The left panel reports wall-clock time per question, and the right panel reports model calls per question. For Nash decoding, stacked bars separate computation before the sequence is fully unmasked from subsequent refinement. All models are evaluated on a single NVIDIA B200 GPU.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value
Prompts, tokens per prompt, continuation slots T
500 , 512 , 64
Action set (vocabulary minus four special IDs)
50,364
Trajectories reaching an exact equilibrium
494 ( 98.8% )
Trajectories stopped at a repeated state
6 ( 1.2% )
Already at equilibrium when construction ends
118 ( 23.6% )
Refinement updates per text (mean / median / max)
2.60 / 2 / 36
Appendix
Table 4: Refinement on 500 WikiText-103 continuations with ModernBERT-Large , counted from the completed sequence x(1) . Construction is a separate max-gap order pass of 64 model calls.
CLAPNQ
PubMedQA
CoQA †
Split
dev, answerable
PQA-L test (seed 0)
dev, gold ≥5 tok.
Questions scored
300
500
1,814
Distinct contexts
297
500
459 ( 460 conv.)
References per question
1.62
1
4 ( 2.36 distinct)
Mean length in ModernBERT tokens (words in parentheses)
Context
195.3 ( 170.8 )
287.0 ( 202.3 )
355.2 ( 267.0 )
Appendix
Table 5: Scale and length of the evaluated subsets. Token lengths use ModernBERT-Large unless otherwise stated; parenthesized word counts are whitespace-delimited. Reference counts are averages per question, with distinct references counted after normalization. Oracle decoding budgets use the reference answer length in each evaluated model’s tokenizer. † CoQA includes only development questions whose gold answer has at least five ModernBERT tokens (Appendix D.3 ).
CLAPNQ
PubMedQA
CoQA †
Lexical overlap with the context
Gold tokens occurring in the context
97.6%
60.0%
99.1%
Gold content words occurring in the context
97.4%
49.3%
98.9%
Novel bigrams (absent from the context)
10.9%
76.8%
6.5%
Answers with <50% content-word overlap
0.3%
52.0%
0.6%
Arrangement of shared wording
Appendix
Table 6: Lexical overlap and arrangement of passage material. Statistics use normalized text and the most favorable reference for each measure. The sentence-selection diagnostic greedily covers distinct answer content words with a 90% target. Copying baselines score either the full context or its best single sentence against the references. † CoQA uses the filtered subset described in Appendix D.3 .
Gold ≥5 tokens ( n=1,814 )
All ( n=7,983 )
Model
Size
F1
R-L
R-Ls
Calls
Char
F1
R-L
R-Ls
One-shot masked LMs (not fine-tuned)
RoBERTa-Large ( Liu et al., 2019 )
355M
28.00
31.45
31.46
1.0
21
47.46
49.29
49.29
ModernBERT-Large ( Warner et al., 2025 )
395M
39.54
42.92
42.93
1.0
21
58.42
60.17
60.17
Ettin-400m ( Weller et al., 2025 )
396M
43.87
46.22
46.24
1.0
21
66.00
66.83
66.83
Autoregressive decoders
Appendix
Table 7: Results on CoQA under the GPT-2 paper prompt . The left block scores the 1,814 questions whose gold answer has ≥5 ModernBERT tokens; the right block scores all 7,983 CoQA development turns , including the yes/no/unknown answers. The protocol is identical throughout and identical to Table 1 : oracle answer length in each model’s own tokenizer, structural-token bans only, greedy decoding, fp32. Calls is the mean number of forward passes per example; for autoregressive models it equals the oracle token budget B . Char is the mean character length of the scored answer. Calls and Chars are reported for the ≥5 cohort.
Model
Size
F1
R-L
R-Ls
Calls
Char
One-shot masked LMs (not fine-tuned)
RoBERTa-Large ( Liu et al., 2019 )
355M
7.99
8.21
8.74
1.0
176
ModernBERT-Large ( Warner et al., 2025 )
395M
5.93
7.95
8.34
1.0
93
Ettin-400m ( Weller et al., 2025 )
396M
4.03
4.55
4.76
1.0
63
Autoregressive decoders
SmolLM2-135M ( Allal et al., 2025 )
135M
22.45
20.23
20.93
48.8
236
Appendix
Table 8: Full results on PubMedQA.
Model
Size
F1
R-L
R-Ls
Calls
Char
One-shot masked LMs (not fine-tuned)
RoBERTa-Large ( Liu et al., 2019 )
355M
5.94
10.38
10.59
1.0
195
ModernBERT-Large ( Warner et al., 2025 )
395M
7.03
12.05
12.66
1.0
142
Ettin-400m ( Weller et al., 2025 )
396M
6.97
11.78
12.41
1.0
117
Autoregressive decoders
GPT-2 Small ( Radford et al., 2019 )
124M
29.31
29.34
29.30
63.1
282
Appendix
Table 9: Full results on CLAPNQ.
Figure 5: Answer F1 (green, right scale) and the Nash gap G(x(k)) (violet, left scale) against the update index k from the all- [MASK] canvas, for ModernBERT-Large with all-mask construction.
Model
Prompt
F1
R-L
R-Ls
Autoregressive decoding
Falcon-7B
Question: / Answer:
32.87
29.80
32.61
Falcon-7B
Soft-instruction
34.39
32.91
34.50
Nash decoding
ModernBERT-Large
Question: / Answer:
37.00
33.25
36.39
ModernBERT-Large
Soft-instruction
38.04
34.79
37.42
Appendix
Table 10: Prompt-format ablation on CLAPNQ.
Calls
Dataset
Mean T
Diffusion
Nash gap
CoQA ( 1,814 )
7.5
8
53
PubMedQA ( 500 )
48.3
48
1,604
CLAPNQ ( 300 )
63.8
64
3,076
Appendix
Table 11: Cost and convergence of the three coordinate rules. Calls is the mean number of forward passes per example. The diffusion variants always terminate after T commitments.
Figure 6: Wall-clock seconds and model calls per question for the same five systems on each dataset, measured on a single B200 in float32 with TF32 disabled and one evaluation process running at a time. The Nash bars are split at the update that first fills the last [MASK] : the light segment includes all computation through that update, and the dark segment includes subsequent refinement. For all-mask decoding, the light segment can include revisions of previously filled positions. Horizontal scales differ between panels. CoQA and PubMedQA use their full evaluation sets. CLAPNQ uses a fixed 60 -item subsample, consisting of every fifth item and shared by all five systems. Its mean reference answer length is 49.0 words, compared with 51.7 for the full set; mean L2R cost is 249 calls, compared with 248 on all 300 items.
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.
Scott Geng, Yufei Zhang, Joseph Lee +3
University of Washington · Meta Superintelligence Labs
Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them from the same text, and the resulting pairs are used to fine-tune, distill, or compress knowledge into another model. We show that this generation step is not neutral preprocessing. It is an implicit policy that both selects which evidence becomes training signal and decides how that evidence is answered, and it is fragile at both stages. When choosing what to ask, generators do not scan a document uniformly. Coverage saturates early and concentrates on salient spans, diverse prompts converge on the same regions, and what looks question-worthy is driven by local presentation. As a result, salient artifacts such as poorly cleaned markup can hijack question generation across model families and scales. When answering, the model that produces the supervision tends to obey instruction-like passages embedded in the text. This compliance depends on the intent and surface form of the passage rather than its strictness, and is worst under task conflict, where larger models comply more often. These failure modes arise from choices made during QA generation, so they can be reduced without changing the training loop. Tying each question to a fixed target reduces biased selection, and filtering instruction-like spans before answering lowers mean injection compliance from 88% to 13% in our evaluation while retaining nearly all clean text.
Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for what follows. Using controlled corruption, we decouple these two roles and find a reversal: making a noisy token easier to predict reduces its damage as a target but increases it as context. The same decoupling helps explain text generated by language models: generation selects each token by its fit to the prefix, while its role as context is never tested against an independently determined continuation, because that continuation is generated to fit it. At known corrupted positions, acting through the context can reduce damage that removing the token's own loss does not. Understanding and controlling what a model learns from a token therefore requires decoupling its roles.
Suqin Yuan, Runqi Lin, Kevin Qinghong Lin +4
University of Sydney · University of Oxford · Southeast University