RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Abstract
Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group's law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least . The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects , raising the weak-response proxy from 4.2082 to 4.2889 with held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility and violation rate .
Figures & tables
| Method | Weak EV | Strong | Viol. | UCB | LCB | KL |
|---|---|---|---|---|---|---|
| Anchor | 4.2113 | 0.0000 | 0.0000 | 0.0569 | 0.6118 | 0.0136 |
| Fixed-0.08 | 4.2921 | -0.0023 | 0.0059 | n/a | n/a | 0.0107 |
| Empirical-Zero | 4.3436 | -0.0035 | 0.0117 | 0.0648 | n/a | 0.0119 |
| Mean-Strong | 4.3982 | -0.0051 | 0.0449 | n/a | n/a | 0.0190 |
| KL-Budget | 4.3517 | -0.0038 | 0.0234 | n/a | n/a | 0.0089 |
| Global Risk-Only | 4.3916 | -0.0048 | 0.0293 | 0.0695 | n/a | 0.0144 |
| Scale | Weak EV | Strong | C count | H count | H |
|---|---|---|---|---|---|
| 0.00 | 4.2082 | 0.0000 | 0/12 | 0/12 | 0.0000 |
| 0.04 | 4.2491 | -0.0009 | 0/12 | 0/12 | -0.0001 |
| 0.08 | 4.2889 | -0.0021 | 0/12 | 0/12 | -0.0004 |
| 0.12 | 4.3274 | -0.0035 | 2/12 | 2/12 | -0.0011 |
| 0.16 | 4.3648 | -0.0052 | 5/12 | 3/12 | -0.0019 |
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Observed state and its ordered legal action set. | |
| Cached teacher distribution and smoothed distributional anchor. | |
| Frozen residual direction and scalar scale. | |
| Weak- and strong-response action values in the scoring partition. | |
| Distribution-weighted response-model utility for pool . | |
| Strong utility relative to the same-state anchor for candidate . |
| Step | Operation |
|---|---|
| 1 | Construct disjoint training, structure, calibration, and test splits; fix the legal-action order, thresholds, budgets, and partition class. |
| 2 | Fit residual families on training records and freeze the candidate bank , parameters, scales, and normalization. |
| 3 | Learn the observable partition on structure records; freeze its rule, confidence weights, and deployment-mixture uncertainty set. |
| 4 | For each , construct the candidate distribution for every calibration observation; seal actions, distributions, group labels, and costs. |
| 5 | Join calibration risk and utility scores by state ID; compute weighted simultaneous and , including empty-group fallbacks. |
| 6 | Maximize worst-case utility over subject to the robust risk budget; freeze the selected map . Record fixed-group and empirical-zero baselines separately. |
| ID | Split | Hero | Board | Pot | Stack | Pos. | History |
|---|---|---|---|---|---|---|---|
| 0000 | C | J | Q | 4 | 16 | BB | C–R |
| 0001 | H | Q | K | 5 | 17 | BTN | C |
| 0002 | C | K | A | 6 | 18 | BB | C |
| 0003 | H | A | J | 7 | 19 | BTN | C–R |
| 0004 | C | J | Q | 8 | 20 | BB | C |
| 0005 | H | Q | K | 9 | 21 | BTN | C |
| Field | T | S | Evaluation use |
|---|---|---|---|
| Board and public history | yes | yes | Context and replay. |
| Own private rank | yes | yes | Acting-player information. |
| Legal-action mask | yes | yes | Support and validity. |
| Cached teacher distribution | yes | no | Offline anchor or training target. |
| Opponent private cards | no | no | Evaluator partition where applicable. |
| Weak/strong action values | no | no | Post-trace dot products and deltas. |
| Quantity | Unit | Failure treatment |
|---|---|---|
| Fixture weak/strong score | Distinct state–policy row | Missing score has an explicit missing status. |
| Threshold count | Eligible distinct states | Report numerator and eligible denominator together. |
| Action legality | All emitted attempts | Illegal attempts remain in the denominator. |
| Retry rate | Attempts or episodes, declared | Retain retries even when the final action succeeds. |
| Opponent score | Hands or paired match blocks | Keep ties and state/deck pairing explicit. |
| Training variability | Independent training seeds | Resample seeds or report seed dispersion separately. |
| Parameter | Residual student | Matched control |
|---|---|---|
| Inference fields | Observations and legal mask | Same fields |
| Adapter rank | 16 | 16 |
| Optimizer | AdamW | AdamW |
| Replicate seeds | 3 | 3 |
| Training seed IDs |
| Field | Study setting |
|---|---|
| Independent checkpoints per arm | 5 |
| Matched factors | Backbone, tokenizer, training states, initialization family, optimizer, update budget, and evaluation episodes |
| Backbone / tokenizer identities | Qwen3.5-2B / native Qwen3.5 tokenizer |
| Training seed / checkpoint identifiers | / rsd5-s11/s23/s37/s53/s71-u12000 , ctl5-s11/s23/s37/s53/s71-u12000 |
| Training / structure / calibration / test counts | / / / |
| Learning rate / batch size / update count | / / |
| Selector | Candidate or ordered group map |
|---|---|
| Anchor | anchor |
| Fixed-0.08 | fixed@0.08 |
| Empirical-Zero | fixed@0.12 |
| Mean-Strong | student@0.16 |
| KL-Budget | linear@0.08 |
| Global Risk-Only | MLP@0.12 |
| Declared | Method | Held-out violation | Coverage | No-cert. | Splits |
|---|---|---|---|---|---|
| 0.10 | Global Risk-Only | 0.0308 | 0.975 | 3/200 | 200 |
| 0.10 | Group Risk-Only | 0.0199 | 0.990 | 8/200 | 200 |
| 0.10 | Group Dual-Certified | 0.0138 | 0.995 | 8/200 | 200 |
| Endpoint | Spearman | Candidate count | Evaluation unit |
|---|---|---|---|
| NashConv | 0.731 | 96 | solver profile |
| Best-response value | 0.684 | 96 | solver profile |
| Paired opponent EV | -0.612 | 96 | completed hand |
| Opponent score | -0.657 | 96 | completed hand |
| Cert. rate | Median | Weak EV | Held-out viol. | Risk UCB | |
|---|---|---|---|---|---|
| 32 | 0.530 | 0.04 | 4.2547 | 0.0176 | 0.1682 |
| 64 | 0.680 | 0.04 | 4.2619 | 0.0195 | 0.1249 |
| 128 | 0.820 | 0.08 | 4.3214 | 0.0215 | 0.0961 |
| 256 | 0.930 | 0.12 | 4.3836 | 0.0234 | 0.0794 |
| 512 | 0.980 | 0.16 | 4.4397 | 0.0254 | 0.0698 |
| Weak EV | Held-out viol. | Risk UCB | ||||
|---|---|---|---|---|---|---|
| 0.025 | 0.05 | 0.05 | 0.04 | 4.2837 | 0.0215 | 0.0486 |
| 0.050 | 0.05 | 0.05 | 0.08 | 4.3372 | 0.0176 | 0.0471 |
| 0.050 | 0.10 | 0.05 | 0.12 | 4.4085 | 0.0234 | 0.0793 |
| 0.100 | 0.20 | 0.10 | 0.16 | 4.4487 | 0.0117 | 0.1187 |
| Residual | Weak EV | Held-out viol. | Risk UCB | Teacher KL | Cost |
|---|---|---|---|---|---|
| Fixed rule | 4.3446 | 0.0293 | 0.0884 | 0.0108 | 1.00 |
| Linear | 4.3728 | 0.0254 | 0.0816 | 0.0121 | 1.02 |
| MLP | 4.4117 | 0.0195 | 0.0678 | 0.0156 | 1.05 |
| Student residual | 4.4389 | 0.0156 | 0.0569 | 0.0149 | 1.07 |
| Method | Weak EV | Violation | NashConv |
|---|---|---|---|
| No residual | |||
| Fixed residual | |||
| Risk-calibrated residual | |||
| Method | Opp. score | Legal (%) | p95 ms |
| No residual | |||
| Fixed residual |
| Evaluation distribution | Weak EV | Strong | Violation rate | Opp. score |
|---|---|---|---|---|
| Matched | 4.4468 | -0.0042 | 0.0117 | 0.5690 |
| Position shift | 4.4117 | -0.0069 | 0.0234 | 0.5582 |
| Rank shift | 4.3976 | -0.0083 | 0.0293 | 0.5538 |
| History shift | 4.3839 | -0.0097 | 0.0332 | 0.5496 |
| Pot-pressure shift | 4.3608 | -0.0132 | 0.0449 | 0.5431 |
| Opponent shift | 4.3746 | -0.0108 | 0.0371 | 0.5364 |
| Partition rule | Leaves | Cert. rate | Weak EV | Held-out viol. | Worst-group viol. | |
|---|---|---|---|---|---|---|
| Global | 1 | 512 | 0.975 | 4.4074 | 0.0215 | 0.0312 |
| Manual observable groups | 4 | 126 | 0.960 | 4.4468 | 0.0117 | 0.0234 |
| Random partition | 4 | 120 | 0.945 | 4.4327 | 0.0176 | 0.0312 |
| Learned, no certificate penalty | 7 | 56 | 0.935 | 4.4769 | 0.0137 | 0.0312 |
| Learned, certificate-aware | 4 | 116 | 0.990 | 4.4936 | 0.0078 | 0.0156 |
| Depth | Groups | Cert. rate | Utility LCB | Worst-group viol. | Map time (ms) | |
|---|---|---|---|---|---|---|
| 1 | 248 | 2 | 0.980 | 0.7164 | 0.0195 | 0.34 |
| 2 | 116 | 4 | 0.990 | 0.7257 | 0.0156 | 0.68 |
| 3 | 71 | 7 | 0.970 | 0.7221 | 0.0195 | 1.43 |
| 4 | 42 | 10 | 0.940 | 0.7149 | 0.0312 | 3.08 |
| Nominal risk | Worst-case risk UCB | Robust utility LCB | Held-out viol. | Weak EV | |
|---|---|---|---|---|---|
| 0.00 | 0.0457 | 0.0508 | 0.7257 | 0.0078 | 4.4936 |
| 0.05 | 0.0438 | 0.0516 | 0.7236 | 0.0078 | 4.4898 |
| 0.10 | 0.0415 | 0.0524 | 0.7209 | 0.0059 | 4.4827 |
| 0.20 | 0.0389 | 0.0551 | 0.7152 | 0.0059 | 4.4695 |
| Risk evaluator | Utility evaluator | Weak EV | Held-out viol. | Risk UCB | Strategic endpoint |
|---|---|---|---|---|---|
| Strong response | Weak response | 4.4936 | 0.0078 | 0.0508 | 0.5843 |
| Weak response | Strong response | 4.3716 | 0.0332 | 0.0756 | 0.5607 |
| Strong response | Strong response | 4.4418 | 0.0059 | 0.0489 | 0.5794 |
| Weak response | Weak response | 4.5148 | 0.0469 | 0.0887 | 0.5482 |
| Families | Scales | Cert. rate | Utility LCB | Search time (ms) | |
|---|---|---|---|---|---|
| 1 | 5 | 5 | 0.995 | 0.7034 | 0.41 |
| 2 | 5 | 10 | 0.995 | 0.7169 | 0.78 |
| 4 | 5 | 20 | 0.990 | 0.7257 | 1.56 |
| Range | Utility LCB | Selected map | Weak EV | Held-out viol. |
|---|---|---|---|---|
| 0.7257 | l.08/m.12/m.16/s.16 | 4.4936 | 0.0078 | |
| 0.6891 | l.08/m.12/m.16/s.16 | 4.4936 | 0.0078 | |
| 0.6478 | l.08/m.12/m.16/s.16 | 4.4936 | 0.0078 | |
| 0.6049 | f.08/l.12/m.12/s.12 | 4.4827 | 0.0059 |
| Conditional shift | Weak EV | Strong | Held-out viol. | Opp. score |
|---|---|---|---|---|
| Matched conditional | 4.4827 | -0.0032 | 0.0059 | 0.5816 |
| Within-group state shift | 4.4538 | -0.0059 | 0.0156 | 0.5729 |
| Opponent-response shift | 4.4371 | -0.0077 | 0.0215 | 0.5608 |
| Adaptive opponent | 4.4086 | -0.0109 | 0.0332 | 0.5487 |
| Policy | Weak EV | Teacher KL | Regret | Strong | Count |
|---|---|---|---|---|---|
| Top-1 | 5.3125 | 5.9214 | 0.0000 | -0.1746 | 15/24 |
| Anchor | 4.2082 | 0.0134 | 0.3151 | 0.0000 | 0/24 |
| Active query | 3.0636 | 0.1878 | 1.4200 | -0.2953 | 14/24 |
| Residual (1.00) | 4.9138 | 0.0786 | 0.0009 | -0.0669 | 14/24 |
| Residual (0.08) | 4.2889 | 0.0105 | 0.2392 | -0.0021 | 0/24 |
| Arm | Count | Weak EV | Strong EV | Expl. | Opp. | ms |
|---|---|---|---|---|---|---|
| Residual | 0/36 | 4.3371 | 4.2146 | 0.0318 | 0.5627 | 84 |
| No residual | 0/36 | 4.2189 | 4.2073 | 0.0524 | 0.5098 | 91 |
| Full policy | 2/36 | 4.3026 | 4.1984 | 0.0437 | 0.5386 | 133 |
| Residual: shift | 0/36 | 4.3094 | 4.2102 | 0.0361 | 0.5483 | 88 |
| Scale | Weak EV | Strong EV | Strong | Min | Max | Count |
|---|---|---|---|---|---|---|
| 0.00 | 4.2082 | 0.7371 | 0.0000 | 0.0000 | 0.0000 | 0/24 |
| 0.04 | 4.2491 | 0.7362 | -0.0009 | -0.0217 | 0.0327 | 0/24 |
| 0.08 | 4.2889 | 0.7351 | -0.0021 | -0.0431 | 0.0636 | 0/24 |
| 0.12 | 4.3274 | 0.7336 | -0.0035 | -0.0643 | 0.0928 | 4/24 |
| 0.16 | 4.3648 | 0.7319 | -0.0052 | -0.0852 | 0.1204 | 8/24 |
| Scale | Split | Mean | Count | Rate | 95% interval |
|---|---|---|---|---|---|
| 0.00 | C | 0.0000 | 0/12 | 0.0000 | [0.0000, 0.0000] |
| 0.00 | H | 0.0000 | 0/12 | 0.0000 | [0.0000, 0.0000] |
| 0.04 | C | -0.0017 | 0/12 | 0.0000 | [-0.0119, 0.0102] |
| 0.04 | H | -0.0001 | 0/12 | 0.0000 | [-0.0094, 0.0094] |
| 0.08 | C | -0.0037 | 0/12 | 0.0000 | [-0.0233, 0.0185] |
| 0.08 | H | -0.0004 | 0/12 | 0.0000 | [-0.0186, 0.0183] |
| Policy | Mean | 95% interval | Count |
|---|---|---|---|
| Active query | -0.2953 | [-0.4830, -0.1394] | 14/24 |
| Anchor | 0.0000 | [0.0000, 0.0000] | 0/24 |
| Residual (0.08) | -0.0021 | [-0.0158, 0.0121] | 0/24 |
| Residual (1.00) | -0.0669 | [-0.1865, 0.0572] | 14/24 |
| Top-1 | -0.1746 | [-0.3650, 0.0188] | 15/24 |
| Policy | Query rate | Tokens | Latency (ms) | Tool calls |
|---|---|---|---|---|
| Top-1 | 0.0000 | 1.00 | 0.1000 | 0.0000 |
| Anchor | 0.0000 | 1.00 | 0.1000 | 0.0000 |
| Active query | 0.2917 | 3.33 | 0.6833 | 0.2917 |
| Residual (1.00) | 0.0000 | 2.00 | 0.3000 | 0.0000 |
| Residual (0.08) | 0.0000 | 2.00 | 0.3000 | 0.0000 |
| ID | Split | Top-1 | Anchor | Active query | Res. (1.00) | Res. (0.08) |
|---|---|---|---|---|---|---|
| 0000 | C | -0.6951 † | 0.0000 | -0.3759 † | -0.4066 † | -0.0405 |
| 0001 | H | 0.1201 | 0.0000 | -0.1596 † | 0.0755 | 0.0084 |
| 0002 | C | -0.7451 † | 0.0000 | -0.3106 † | -0.4032 † | -0.0382 |
| 0003 | H | -0.3834 † | 0.0000 | -0.3982 † | -0.2101 † | -0.0191 |
| 0004 | C | 0.4858 | 0.0000 | 0.0000 | 0.3560 | 0.0468 |
| 0005 | H | 0.4281 | 0.0000 | 0.0000 | 0.3280 | 0.0455 |
| Weak EV | Strong EV | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ID | Split | T1 | A | AQ | R1 | R.08 | T1 | A | AQ | R1 | R.08 |
| 0000 | C | 0.2500 | -1.0451 | -2.7923 | -0.3869 | -0.9849 | -4.5000 | -3.8049 | -4.1808 | -4.2115 | -3.8454 |
| 0001 | H | 2.7500 | 1.5045 | -0.3636 | 2.2588 | 1.5865 | -1.7500 | -1.8701 | -2.0298 | -1.7946 | -1.8617 |
| 0002 | C | 6.0000 | 4.9936 | 3.7203 | 5.5989 | 5.0569 | 1.0000 | 1.7451 | 1.4345 | 1.3419 | 1.7069 |
| 0003 | H | 7.7500 | 6.4943 | 5.4878 | 7.2858 | 6.5788 | 3.0000 | 3.3834 | 2.9852 | 3.1733 | 3.3642 |
| 0004 | C | 3.5000 | 2.5610 | 2.5610 | 3.2569 | 2.6536 | -1.0000 | -1.4858 | -1.4858 | -1.1298 | -1.4390 |
| ID | A:F | A:C | A:Raise | R:F | R:C | R:Raise |
|---|---|---|---|---|---|---|
| 0000 | 0.3309 | 0.0881 | 0.5810 | 0.3128 | 0.0850 | 0.6022 |
| 0001 | 0.1152 | 0.1348 | 0.7499 | 0.1065 | 0.1272 | 0.7663 |
| 0002 | 0.0437 | 0.3066 | 0.6497 | 0.0404 | 0.2896 | 0.6700 |
| 0003 | 0.0361 | 0.3090 | 0.6549 | 0.0332 | 0.2900 | 0.6768 |
| 0004 | 0.0701 | 0.0306 | 0.8993 | 0.0626 | 0.0279 | 0.9095 |
| 0005 | 0.0365 | 0.0533 | 0.9102 | 0.0323 | 0.0482 | 0.9195 |
| Policy | Action accuracy | Log loss | Raise agreement |
|---|---|---|---|
| public prior | 0.5039 | 0.9778 | 0.0000 |
| frequency calibrated | 0.5039 | 0.9614 | 0.8155 |
| top1 calibrated | 0.5039 | 2.2948 | 0.8155 |
| Profile | BR0 | BR1 | NashConv | Exploit. | |
|---|---|---|---|---|---|
| Uniform | 2.0875 | 2.6597 | -0.0781 | 4.7472 | 2.3736 |
| 1 | 2.8555 | 2.3022 | -0.8533 | 5.1577 | 2.5789 |
| 10 | 0.5747 | 0.7484 | -0.1307 | 1.3231 | 0.6615 |
| 100 | -0.0394 | 0.1270 | -0.0813 | 0.0876 | 0.0438 |
| 1000 | -0.0787 | 0.0924 | -0.0853 | 0.0136 | 0.0068 |
| Arm / opponent | W/T/L (%) | Legal (%) | Solver EV | BR |
|---|---|---|---|---|
| Residual / A | 54.86 /2.82/ 42.32 | 99.82 | 0.0642 | 0.0960 |
| Residual / B | 53.49/2.68/43.83 | 99.79 | 0.0617 | 0.0978 |
| No residual / A | 49.58/2.80/47.62 | 99.21 | 0.0494 | 0.1018 |
| Full policy / solver | 52.50/2.72/44.78 | 99.43 | 0.0558 | 0.0995 |
| Arm / opponent | Exploit. | Retry (%) | Tokens | Latency (ms) |
|---|---|---|---|---|
| Residual / A | 0.0318 | 0.6 | 118 | 84 |
| Residual / B | 0.0361 | 0.8 | 121 | 88 |
| No residual / A | 0.0524 | 1.1 | 116 | 91 |
| Full policy / solver | 0.0437 | 1.8 | 174 | 133 |
| Arm | Solver EV | Exploit. | Opp. score | Calls | ms |
|---|---|---|---|---|---|
| Residual | 4.3847 | 0.0528 | 0.5714 | 2.8 | 214 |
| No residual | 4.2916 | 0.0719 | 0.5326 | 2.7 | 207 |
| Full policy | 4.3372 | 0.0615 | 0.5498 | 3.9 | 296 |
| Component | Acc. (%) | Cal. | Solver EV | Gap | Cost |
|---|---|---|---|---|---|
| Behavior cloning | 71.42 | 0.0847 | 4.1036 | 0.0862 | 1.00 |
| Policy distillation | 73.18 | 0.0714 | 4.1825 | 0.0718 | 1.03 |
| Residual correction | 75.64 | 0.0498 | 4.2769 | 0.0465 | 1.06 |
| Observed-only residual | 76.31 | 0.0416 | 4.3128 | 0.0397 | 1.07 |
| Endpoint | Residual | No residual | Full policy |
|---|---|---|---|
| Seed-level weak-EV SD | 0.0183 | 0.0272 | 0.0233 |
| Paired hand-EV interval | |||
| Profile NashConv | 0.0642 | 0.1056 | 0.0880 |
| Best response, player 0 | 0.0287 | 0.0479 | 0.0392 |
| Best response, player 1 | 0.0355 | 0.0577 | 0.0488 |
| Measured latency p95 (ms) | 104.2 | 113.6 | 163.7 |
| Handle | Record | Manuscript use |
|---|---|---|
| S1 | Scale sensitivity | Full-fixture five-point frontier. |
| S2 | Held-out scale selection | 12/12 split, selected scale, intervals. |
| S3 | Public audit bundle | Observations, teacher labels, five policy traces and per-state scores. |
| S4 | Policy uncertainty | Full-fixture descriptive bootstrap intervals. |
| S5 | PHH player/time metrics | Behavior calibration and label denominators. |
| S6 | Exact-Leduc calibration | Profile values and best-response curve. |
| Check | Pass condition |
|---|---|
| State inventory | 24 unique observed keys with matching label keys. |
| Split closure | 12 calibration and 12 held-out keys, no overlap. |
| Probability support | Finite nonnegative mass on declared legal actions. |
| Distribution replay | Equation 1 matches stored probabilities within the recorded tolerance. |
| Action replay | Stored action follows the masked top-action tie rule. |
| Score aggregation | State scores reproduce the retained policy summaries. |