Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation
Authors: Alireza Jafari, Arman Adibi, Mohammad Ghavamzadeh, Hadi Daneshmand
Organizations: Department of Computer Science University of Virginia Charlottesville, VA, USA · School of Computer and Cyber Sciences Augusta University Augusta, GA, USA · Qualcomm AI Research San Diego, CA, USA
Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an ε-Nash equilibrium in O(1/ε) time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to 18× larger, without any fine-tuning or retraining, at the cost of additional test-time computation.
Figures & tables
Figure 1: Revision Game. Each token position is a player whose choice is evaluated given all context tokens. Autoregressive decoding leaves earlier tokens fixed, whereas the tokens are revised in the revision game. We used GPT-2 small for conditional probability estimates. At a Nash equilibrium, no player can increase its conditional probability by changing its token. Details are reported in Appendix A .
Figure 2: Convergence of Nash decoding on WikiText-103. (a) Running minimum of the Nash gap, compared with a 1/k reference curve. The reference indicates a rate of decay, aligned with Theorem 1 . Both axes are log-transformed. (b) GPT-2 XL continuation perplexity and ModernBERT pseudo-perplexity. The x-axis is log-transformed.
CoQA
PubMedQA
CLAPNQ
Model
Size
F1
R-L
R-Ls
F1
R-L
R-Ls
F1
R-L
R-Ls
One-shot masked LMs (not fine-tuned)
RoBERTa-L ( Liu et al., 2019 )
355M
28.00
31.45
31.46
7.99
8.21
8.74
5.94
10.38
10.59
ModernBERT-L ( Warner et al., 2025 )
395M
39.54
42.92
42.93
5.93
7.95
8.34
7.03
12.05
12.66
Ettin-400m ( Weller et al., 2025 )
396M
43.87
46.22
46.24
4.03
4.55
4.76
6.97
11.78
12.41
Autoregressive models
Table 1: Results across three question-answering benchmarks under a shared evaluation protocol. Size denotes the number of model parameters; for Nash decoding, it denotes the masked model size. Best results are bolded and second-best results are underlined . The complete benchmark results are reported in Appendix F .
Figure 3: F1 score and Nash gap during Nash decoding iterations on CLAPNQ.
Model
Size
F1
R-L
R-Ls
Left-to-right initialization
mmBERT-base
308M
4.22
6.11
5.38
RoBERTa-Large
355M
22.23
19.79
20.77
ModernBERT-Large
395M
21.91
19.88
19.99
Ettin-400m
396M
17.41
17.29
17.36
All-mask initialization
Table 2: Nash Equilibria Basins on PubMedQA.
CoQA
PubMedQA
CLAPNQ
Decoding rule
F1
R-L
R-Ls
F1
R-L
R-Ls
F1
R-L
R-Ls
Uniform-ordered
38.65
40.05
40.08
30.50
23.51
27.14
56.15
50.08
55.00
Confidence-ordered
41.85
42.21
42.28
31.74
24.64
28.50
58.84
54.30
58.11
Max Nash gap
43.07
43.22
43.31
32.11
24.80
28.75
59.16
54.67
58.37
Table 3: Nash decoding versus single-token masked-diffusion decoding using the same masked model. All rows use the instruction-tuned ModernBERT-chat with its corresponding prompt; details in Appendix H .
Figure 4: Runtime comparison of Nash and autoregressive decoding on CoQA. The left panel reports wall-clock time per question, and the right panel reports model calls per question. For Nash decoding, stacked bars separate computation before the sequence is fully unmasked from subsequent refinement. All models are evaluated on a single NVIDIA B200 GPU.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value
Prompts, tokens per prompt, continuation slots T
500 , 512 , 64
Action set (vocabulary minus four special IDs)
50,364
Trajectories reaching an exact equilibrium
494 ( 98.8% )
Trajectories stopped at a repeated state
6 ( 1.2% )
Already at equilibrium when construction ends
118 ( 23.6% )
Refinement updates per text (mean / median / max)
2.60 / 2 / 36
Appendix
Table 4: Refinement on 500 WikiText-103 continuations with ModernBERT-Large , counted from the completed sequence x(1) . Construction is a separate max-gap order pass of 64 model calls.
CLAPNQ
PubMedQA
CoQA †
Split
dev, answerable
PQA-L test (seed 0)
dev, gold ≥5 tok.
Questions scored
300
500
1,814
Distinct contexts
297
500
459 ( 460 conv.)
References per question
1.62
1
4 ( 2.36 distinct)
Mean length in ModernBERT tokens (words in parentheses)
Context
195.3 ( 170.8 )
287.0 ( 202.3 )
355.2 ( 267.0 )
Appendix
Table 5: Scale and length of the evaluated subsets. Token lengths use ModernBERT-Large unless otherwise stated; parenthesized word counts are whitespace-delimited. Reference counts are averages per question, with distinct references counted after normalization. Oracle decoding budgets use the reference answer length in each evaluated model’s tokenizer. † CoQA includes only development questions whose gold answer has at least five ModernBERT tokens (Appendix D.3 ).
CLAPNQ
PubMedQA
CoQA †
Lexical overlap with the context
Gold tokens occurring in the context
97.6%
60.0%
99.1%
Gold content words occurring in the context
97.4%
49.3%
98.9%
Novel bigrams (absent from the context)
10.9%
76.8%
6.5%
Answers with <50% content-word overlap
0.3%
52.0%
0.6%
Arrangement of shared wording
Appendix
Table 6: Lexical overlap and arrangement of passage material. Statistics use normalized text and the most favorable reference for each measure. The sentence-selection diagnostic greedily covers distinct answer content words with a 90% target. Copying baselines score either the full context or its best single sentence against the references. † CoQA uses the filtered subset described in Appendix D.3 .
Gold ≥5 tokens ( n=1,814 )
All ( n=7,983 )
Model
Size
F1
R-L
R-Ls
Calls
Char
F1
R-L
R-Ls
One-shot masked LMs (not fine-tuned)
RoBERTa-Large ( Liu et al., 2019 )
355M
28.00
31.45
31.46
1.0
21
47.46
49.29
49.29
ModernBERT-Large ( Warner et al., 2025 )
395M
39.54
42.92
42.93
1.0
21
58.42
60.17
60.17
Ettin-400m ( Weller et al., 2025 )
396M
43.87
46.22
46.24
1.0
21
66.00
66.83
66.83
Autoregressive decoders
Appendix
Table 7: Results on CoQA under the GPT-2 paper prompt . The left block scores the 1,814 questions whose gold answer has ≥5 ModernBERT tokens; the right block scores all 7,983 CoQA development turns , including the yes/no/unknown answers. The protocol is identical throughout and identical to Table 1 : oracle answer length in each model’s own tokenizer, structural-token bans only, greedy decoding, fp32. Calls is the mean number of forward passes per example; for autoregressive models it equals the oracle token budget B . Char is the mean character length of the scored answer. Calls and Chars are reported for the ≥5 cohort.
Model
Size
F1
R-L
R-Ls
Calls
Char
One-shot masked LMs (not fine-tuned)
RoBERTa-Large ( Liu et al., 2019 )
355M
7.99
8.21
8.74
1.0
176
ModernBERT-Large ( Warner et al., 2025 )
395M
5.93
7.95
8.34
1.0
93
Ettin-400m ( Weller et al., 2025 )
396M
4.03
4.55
4.76
1.0
63
Autoregressive decoders
SmolLM2-135M ( Allal et al., 2025 )
135M
22.45
20.23
20.93
48.8
236
Appendix
Table 8: Full results on PubMedQA.
Model
Size
F1
R-L
R-Ls
Calls
Char
One-shot masked LMs (not fine-tuned)
RoBERTa-Large ( Liu et al., 2019 )
355M
5.94
10.38
10.59
1.0
195
ModernBERT-Large ( Warner et al., 2025 )
395M
7.03
12.05
12.66
1.0
142
Ettin-400m ( Weller et al., 2025 )
396M
6.97
11.78
12.41
1.0
117
Autoregressive decoders
GPT-2 Small ( Radford et al., 2019 )
124M
29.31
29.34
29.30
63.1
282
Appendix
Table 9: Full results on CLAPNQ.
Figure 5: Answer F1 (green, right scale) and the Nash gap G(x(k)) (violet, left scale) against the update index k from the all- [MASK] canvas, for ModernBERT-Large with all-mask construction.
Model
Prompt
F1
R-L
R-Ls
Autoregressive decoding
Falcon-7B
Question: / Answer:
32.87
29.80
32.61
Falcon-7B
Soft-instruction
34.39
32.91
34.50
Nash decoding
ModernBERT-Large
Question: / Answer:
37.00
33.25
36.39
ModernBERT-Large
Soft-instruction
38.04
34.79
37.42
Appendix
Table 10: Prompt-format ablation on CLAPNQ.
Calls
Dataset
Mean T
Diffusion
Nash gap
CoQA ( 1,814 )
7.5
8
53
PubMedQA ( 500 )
48.3
48
1,604
CLAPNQ ( 300 )
63.8
64
3,076
Appendix
Table 11: Cost and convergence of the three coordinate rules. Calls is the mean number of forward passes per example. The diffusion variants always terminate after T commitments.
Figure 6: Wall-clock seconds and model calls per question for the same five systems on each dataset, measured on a single B200 in float32 with TF32 disabled and one evaluation process running at a time. The Nash bars are split at the update that first fills the last [MASK] : the light segment includes all computation through that update, and the dark segment includes subsequent refinement. For all-mask decoding, the light segment can include revisions of previously filled positions. Horizontal scales differ between panels. CoQA and PubMedQA use their full evaluation sets. CLAPNQ uses a fixed 60 -item subsample, consisting of every fifth item and shared by all five systems. Its mean reference answer length is 49.0 words, compared with 51.7 for the full set; mean L2R cost is 249 calls, compared with 248 on all 300 items.