SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Johns Hopkins University · Microsoft · Fudan University · University of Science and Technology of China
Abstract
A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents' behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupported strategic assumptions, inconsistent opponent estimates, and harmful interference from irrelevant historical interactions. To address these issues, we propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate. SAGE first anchors reasoning to an equilibrium policy that provides a strategically valid prior. It then conditions deviations from this anchor on a soft belief over opponent behavioral tendencies, enabling opponent-specific exploitation. Finally, SAGE distills strategically related interactions into counterfactual hypotheses about previously missing considerations, allowing past experience to recalibrate the model's reasoning. We evaluate SAGE on three repeated imperfect-information games: Leduc Hold'em, Liar's Dice, and Goofspiel, against various opponent types in each game. Compared with reasoning-intensive LLM agents, including Suspicion-Agent, ReTA, Agent-Pro, EMO, and Hypothetical Minds, SAGE achieves up to a 127.6% payoff improvement in Liar's Dice while reducing input and output token usage by up to 80% and 90%, respectively. In direct match-up play, it attains non-negative mean payoff against 5/10, 8/10, and 8/10 evaluated opponents in Leduc Hold'em, Liar's Dice, and Goofspiel, respectively, while using relatively fewer tokens. Code is available at https://github.com/chenzhwsysu57/SAGE.
Figures & tables
| Method | Leduc Hold’em | Liar’s Dice | Goofspiel | |||
|---|---|---|---|---|---|---|
| WR (%) | Payoff | WR (%) | Payoff | WR (%) | Payoff | |
| Vanilla | 45.8 | +0.778 | 60.8 | +0.217 | 52.5 | +1.444 |
| LLM(Eq) | 44.7 | +0.953 | 69.8 | +0.394 | 55.8 | +1.411 |
| LLM(Opp) | 47.0 | +0.928 | 60.0 | +0.200 | 63.9 | +2.050 |
| MCCFR* | 50.8 | +0.914 | 68.3 | +0.367 | 51.7 | +1.081 |
| Suspicion-Agent | 51.4 | +0.717 | 72.7 | +0.222 | 77.6 | +2.247 |
| Opponent | Leduc Hold’em | Liar’s Dice | Goofspiel | |||
|---|---|---|---|---|---|---|
| WR (%) | Payoff | WR (%) | Payoff | WR (%) | Payoff | |
| Vanilla | 63.3 | +1.750 | 68.3 | +0.367 | 55.0 | +1.217 |
| LLM(Eq) | 8.3 | -2.967 | 58.3 | +0.167 | 11.7 | -0.100 |
| LLM(Opp) | 50.0 | -0.467 | 50.0 | 0.000 | 56.7 | +1.533 |
| MCCFR* | 53.3 | +0.500 | 43.3 | -0.133 | 15.0 | -0.033 |
| Suspicion-Agent | 48.3 | -0.250 | 68.3 | +0.367 | 68.6 | +2.647 |
| Removed component | Leduc | Liar’s Dice | Goofspiel |
|---|---|---|---|
| Equilibrium guidance | +0.3111 | +0.2222 | +0.3278 |
| Opponent belief | +0.1250 | +0.0611 | +0.9694 |
| Counterfactual recalibration | +0.0333 | +0.0778 | +0.2722 |
| Method | Game | Strategic | Opponent | Historical | Total | API |
|---|---|---|---|---|---|---|
| context | planning | reasoning | content | tokens | calls | |
| Vanilla | 1855(56.84%) | 1290(39.50%) | 120(3.66%) | 0(0%) | 3265 | 4.21 |
| LLM(eq) | 2936(46.970%) | 2463(39.400%) | 852(13.631%) | 0(0%) | 6251 | 7.25 |
| LLM(op) | 3400(36.95%) | 3001(32.62%) | 2800(30.43%) | 0(0%) | 9201 | 8.22 |
| SAGE | 1771(16.34%) | 1626(15.00%) | 2264(20.89%) | 5177(47.77%) | 10837 | 4.92 |
| AgentPro | 8962(18.58%) | 10649(22.08%) | 6785(14.07%) | 21840(45.28%) | 48236 | 8.27 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric | Ours | Ours (weak) | Ours (code) | Vanilla |
|---|---|---|---|---|
| Strategic reference | MCCFR ( ) | MCCFR ( ) | Search with code | None |
| WR (%) | 55.0 | 50.0 | 51.1 | 45.8 |
| Payoff | 1.139 | 0.825 | 0.842 | 0.778 |
| Game | Observable evidence and comparison |
|---|---|
| Leduc Hold’em | Smoothed check, bet, call, raise, and fold frequencies, conditioned on whether a player faces a bet, together with action entropy and showdown-based selectivity when cards are revealed. These summaries are compared with style signatures estimated from profiling episodes. |
| Liar’s Dice | Bid and challenge frequencies, claim ranks, challenge rates after low and high bids, escalation rates, and first-action challenges. Features are standardized against position-aware prototype statistics before comparison. |
| Goofspiel | Public current-prize–opponent-bid pairs from completed rounds. Each prototype is a smoothed conditional table over prize values and bids, and recent evidence is scored by its log likelihood under these tables. |
| Method | Game | Strategic | Opponent | Historical | Total | API |
|---|---|---|---|---|---|---|
| context | planning | reasoning | content | tokens | calls | |
| Vanilla | 1218(66.85%) | 542(29.74%) | 62(3.42%) | 0(0%) | 1823 | 2.29 |
| LLM(eq) | 1287(47.142%) | 742(27.179%) | 506(18.527%) | 195(7.152%) | 2731 | 2.36 |
| LLM(op) | 1476(38.20%) | 750(19.42%) | 1411(36.52%) | 226(5.86%) | 3863 | 2.74 |
| SAGE | 1183(25.90%) | 891(19.50%) | 914(20.01%) | 1580(34.59%) | 4568 | 2.67 |
| Suspicion | 2846(11.26%) | 6271(24.80%) | 15319(60.58%) | 852(3.37%) | 25288 | 6.66 |
| Method | Game | Strategic | Opponent | Historical | Total | API |
|---|---|---|---|---|---|---|
| context | planning | reasoning | content | tokens | calls | |
| Vanilla | 1562(53.78%) | 1179(40.58%) | 164(5.64%) | 0(0%) | 2904 | 5.00 |
| LLM(eq) | 1375(33.846%) | 2034(50.049%) | 654(16.105%) | 0(0%) | 4063 | 5.00 |
| LLM(op) | 1426(28.79%) | 1613(32.58%) | 1913(38.63%) | 0(0%) | 4951 | 5.00 |
| SAGE | 1466(9.24%) | 2274(14.33%) | 2114(13.32%) | 10015(63.11%) | 15868 | 7.48 |
| AgentPro | 8635(20.40%) | 11345(26.81%) | 5875(13.88%) | 16467(38.91%) | 42321 | 10.13 |
| Game | Shift | Method | Win rate | Net payoff | Pre/Post payoff |
|---|---|---|---|---|---|
| Leduc Hold’em | TAG CallingStation | SAGE | 40.0% | +23 | +4/+19 |
| SAGE(Bayes) | 38.8% | -7 | -24/+17 | ||
| CFR | 42.5% | +10 | -5/+15 | ||
| Vanilla | 41.2% | +4 | -13/+17 | ||
| Suspicion | 42.5% | -16 | -21/+5 | ||
| Liar’s Dice | Bluffer Conservative | SAGE | 71.2% | +34 | +10/+24 |