cs.AISep 27, 2026

RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games

Authors: Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Peng Zhang, Daren Zha, Jun Xiao

Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China

Abstract

Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group's law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least 1−ζrisk−ζutil1-ζ_{risk}-ζ_{util}. The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects α=0.08α=0.08, raising the weak-response proxy from 4.2082 to 4.2889 with 0/120/12 held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility 4.3659±0.01774.3659\pm0.0177 and violation rate 0.0178±0.00570.0178\pm0.0057.

Figures & tables

Appendix figures & tables40 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

    Jun 7, 2026Xiaoli Yu, Jiamiao LiuConformal Risk ControlFinite-Sample Certificates

  2. Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations

    Sep 11, 2026Tong Li, Saunak Kumar Panda, Yisha XiangConditional-Value-At-RiskOffline Reinforcement Learning