OSCC: Certified Observation-Safe Coupling Optimization for Gradient-Noise Control in Imperfect-Information Learning
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Abstract
Coupled rollouts can reduce the noise of counterfactual action comparisons, but two issues prevent standard common-random-number constructions from serving as a general learning primitive in imperfect-information environments. First, an invalid coupling may expose hidden state, synchronize endogenous policy randomness, or misalign chance events after counterfactual histories diverge. Second, in multi-action policy optimization, lower return-contrast variance is not by itself the relevant objective: the optimizer depends on the return covariance matrix after projection through the local policy-gradient geometry. We introduce observation-safe counterfactual coupling (OSCC), a framework that defines an admissible class through marginal preservation, information-state safety, branch-local policy randomness, semantic event alignment, and trace-before-oracle replay. We derive a gradient-aware coupling criterion showing that, for marginal-preserving couplings, policy-gradient noise changes are determined by policy-Jacobian-weighted off-diagonal return covariance. This motivates OSCC-Select, a calibration-only selector that chooses among independent, root-only, continuation-only, and fully coupled rollouts using separate safety and gain certificates. Its gain target combines projected gradient noise with measured physical sampling cost and falls back to independent sampling whenever a simultaneous lower confidence bound does not certify improvement. On 100,000 fixed-root Leduc comparisons, the fully coupled CP-GRPO instantiation reduces return-contrast variance from 41.1158 to 18.1441, a 55.87% reduction, while preserving the declared branch marginals. With three actions, OSCC-Select chooses continuation coupling and attains gradient-noise trace 0.0783 versus 0.0917 for return-variance selection. Increasing calibration from 64 to 2,048 groups raises certification from 0.327 to 0.995.
Figures & tables
| namespace | inputs | shared? | policy visibility and audit |
|---|---|---|---|
| replay/root | run, root, group | yes | no; root hash equality |
| environment | env key, event counter | yes | legal observation only; mask/hash check |
| continuation | env key, public history | yes | through public events; monotone counter |
| policy action | run, group, branch, time | no | action draw only; hidden-card invariance |
| ledger | event, branch, status | no | no effect on rollout; join after trace freeze |
| arm | pairs | reported mean | variance | SD |
|---|---|---|---|---|
| independent | 100000 | 0.01421 | 41.1158 | 6.4123 |
| CP-GRPO paired | 100000 | 0.03487 | 18.1441 | 4.2596 |
| quantity | independent | CP-GRPO paired | paired bootstrap interval |
|---|---|---|---|
| contrast variance | 41.1158 | 18.1441 | |
| variance ratio | 1.0000 | 0.4412 | |
| check–bet covariance | 11.17784 |
| branch | mean difference | 95% CI low | 95% CI high | decision |
|---|---|---|---|---|
| check | 0.02049 | 0.0071 | 0.0342 | equivalent ( ) |
| bet | -0.01115 | -0.0245 | 0.0022 | equivalent ( ) |
Appendix figures & tables40 assets
Supplementary material from the paper’s appendix.
Appendix
| quantity | symbol | unit | role |
|---|---|---|---|
| branch return | reward | contrast endpoint | |
| group mean | reward | centring reference | |
| centred advantage | reward | update input | |
| importance ratio | unitless | policy correction | |
| clipped ratio | unitless | trust-region diagnostic | |
| paired covariance | reward 2 | variance reduction source |
| stage | required record | validation rule |
|---|---|---|
| root construction | root state, seat, legal actions, group key | root hash matches every branch |
| stream derivation | chance and policy namespace, counter start | no counter collision across semantic events |
| online rollout | observations, actions, masks, terminal reason | hidden/oracle fields absent from policy input |
| trace freeze | ordered events, replay digest, failure reason | exact event order regenerates from manifest |
| post-hoc join | terminal labels, solver values, contrast | join occurs after trace hash is sealed |
| branch swap | seat-swapped duplicate and utility sign | contrast reverses under declared convention |
| namespace | seed input | visible online | check |
|---|---|---|---|
| root | run/group key | no | root equality |
| chance | group/event counter | public event only | counter monotonicity |
| policy | group/branch/action counter | action sampling | branch separation |
| opponent | run/opponent ID | opponent response | frozen-opponent match |
| bootstrap | metric/replicate ID | no | interval regeneration |
| method | common state/scenario | forced alternatives | optimization objective | partial-observation contract | safety / fallback | replay / artifact binding |
|---|---|---|---|---|---|---|
| PEGASUS | scenario seed | policy candidates | scalar policy-search comparison | POMDP simulator | no certified fallback | scenario reuse, no hidden-card ledger |
| paired comparisons | common scenario | policy candidates | paired contrast variance | task-dependent | no safety selector | paired statistical comparison |
| TRPO vine | common state and rollout prefix | alternative actions | trust-region update signal | MDP policy state | no coupling fallback | optimizer trace, no game-specific ledger |
| rollout CRN | shared exogenous draws | selected branches | scalar return variance | task-dependent | simulator-dependent | seed or simulator dependent |
| OSCC/CP-GRPO | fixed root and namespaced continuation | legal root actions | gradient noise plus physical cost | information-state only; hidden-card invariance test | C1–C5 and independent fallback | group/stream hashes and replay contract |
| dimension | established precedent | CP-GRPO contract | verification artifact |
|---|---|---|---|
| shared factor | scenario or state prefix | root and namespaced continuation | root hash and event ledger |
| partial observation | task-dependent simulator state | acting player’s information state only | observation hash and hidden-field audit |
| policy randomness | implementation-specific | branch-specific counter namespace | policy-key replay and invariance test |
| comparison unit | candidate or trajectory | legal root-action group | group key and branch-swap check |
| claim boundary | policy comparison or planning return | estimator precision plus provenance | variance identity and manifest hash |
| event | key inputs | policy-visible fields | invariant | failure code |
|---|---|---|---|---|
| root construction | run, root, group, seat | public root projection and own card | root hash is identical across branches | ROOT_DRIFT |
| chance continuation | environment key and event counter | resulting public observation and legal mask | counter is monotone and has no semantic collision | STREAM_GAP |
| policy action | run, group, branch, action counter | information state and legal actions | changing the branch key preserves opponent-card privacy | HIDDEN_LEAK |
| terminal event | branch, terminal reason, consumed counters | terminal public record | unused counters remain reserved after early termination | EARLY_END |
| post-hoc join | trace digest and branch key | no oracle or solver field | solver and clean labels join only after trace freeze | JOIN_ORDER |
| field | scope | visible online? | validation and use |
|---|---|---|---|
| group key | root group | collector only | stable intervention identifier |
| branch key | branch | no | post-hoc branch join |
| root-state hash | root | no | verifies shared starting state |
| observation hash | transition | validator only | checks information-set equality |
| chance-stream counter | event | collector only | detects missing or reused draw |
| policy-stream counter | event | sampler only | separates action randomness |
| screen | primary unit | reported measurement | uncertainty unit | supported decision |
|---|---|---|---|---|
| fixed-root estimator | group | 100,000 valid groups; variance 18.1441 paired versus 41.1158 independent | group bootstrap and covariance identity | estimator precision |
| exact calibration | enumerated iteration | Kuhn 20,000 and Leduc 300 iterations; max exploitability 0.0007619 and 0.0151834 | deterministic rerun | terminal and solver convention |
| frozen-opponent update | branch rollout | 49,152 rollouts per seed; margin SE 0.028117 paired and 0.042755 independent | state-level summary | one-step mechanism diagnostic |
| scalar clipped screen | seed and logical batch | 1,600 finite nonzero gradients, 40 iterations, 32 groups, four epochs | seed-level summary | update wiring and cost accounting |
| policy/solver endpoint | seed and held-out root | 5 seeds; 512 held-out roots/seed; value 0.2136; exploitability 0.1184 | paired seed interval | held-out policy quality |
| check | unit | paired arm | independent arm | pass criterion and interpretation |
|---|---|---|---|---|
| unique group key | replay group | 5120/5120 | 5120/5120 | one logical key per root/action pair |
| observation hash drift | transition | 0/25549 | 0/25549 | same information state under hidden-field perturbation |
| stream counter gap | event | 0/102960 | 0/102880 | monotone chance and policy counters |
| branch-swap sign | group | 5120/5120 | 5120/5120 | contrast sign reverses after branch reordering |
| marginal return shift | root | 0.0018 | 0.0021 | no change in arm marginal target |
| condition | estimator action | update action | reason |
|---|---|---|---|
| observation and mask agree | compute paired contrast | eligible | valid coupling |
| observation hash differs | record diagnostic contrast | ineligible | information-set change |
| stream counter gap | replay from manifest | ineligible | non-reproducible trace |
| early termination with ledger code | retain return and cost | eligible if hashes agree | explicit terminal event |
| oracle field in policy input | quarantine trace | ineligible | post-hoc boundary violated |
| seed | paired variance | independent variance | covariance | ratio | valid groups |
|---|---|---|---|---|---|
| 13 | 18.3921 | 41.4267 | 11.2518 | 0.4439 | 5120 |
| 17 | 17.9846 | 40.8173 | 11.0842 | 0.4406 | 5120 |
| 23 | 18.0557 | 41.1038 | 11.1975 | 0.4393 | 5120 |
| 29 | 18.2814 | 41.3659 | 11.1638 | 0.4419 | 5120 |
| 31 | 18.1948 | 40.8847 | 11.1429 | 0.4451 | 5120 |
| parameter | value/unit | arms | audit output |
|---|---|---|---|
| environment | Leduc, Kuhn transfer | all | environment hash |
| training seeds | five named seeds | all | seed manifest |
| rollout horizon | fixed transitions | all | transition count |
| batch size | fixed groups | all | group denominator |
| optimizer | common schedule | all | optimizer digest |
| clip coefficient | common value | PPO/GRPO/paired | clipped fraction |
| budget | paired arm | control arm | validation |
|---|---|---|---|
| transitions | fixed groups | fixed groups | transition count |
| updates | common schedule | common schedule | optimizer digest |
| solver calls | held-out roots | held-out roots | state enumeration |
| policy calls | ledger count | ledger count | query accounting |
| wall time | matched seconds | matched seconds | hardware class |
| artifact | required fields | validation predicate | manuscript use |
|---|---|---|---|
| run manifest | source path, configuration hash, seed list | hash matches release record | identifies the experiment |
| root ledger | root ID, seat, legal mask, group key | root equality across branches | fixed-root estimand |
| event ledger | namespace, event counter, observation hash | monotone counters and public-state agreement | observation safety |
| summary file | group count, means, covariance, residual | variance identity and zero silent filtering | estimator table |
| figure record | input hash, caption, placement, availability status | source file exists and its hash is recorded | visual evidence |
| seed | episodes | optimizer epochs | logical batches | sampling calls |
| 13 | 1096 | 32 | 8 | 5302 |
| 17 | 1096 | 32 | 8 | 5291 |
| 23 | 1096 | 32 | 8 | 5308 |
| 29 | 1096 | 32 | 8 | 5297 |
| 31 | 1096 | 32 | 8 | 5311 |
| control | unit | result type | reported statistic | interpretation |
|---|---|---|---|---|
| exact solver | fixed enumeration | pass | max exploitability 0.0007619 (Kuhn), 0.0151834 (Leduc) | calibration of terminal returns |
| frozen opponent | one-step update | descriptive | 49,152 rollouts/seed; margin SE 0.028117 paired, 0.042755 independent | post-root mechanism check |
| scalar clipped surrogate | five seeds | mixed paired contrast | 1,600 gradients; selected clipping 0; check delta 0.188967 paired, 0.169950 independent | implementation connectivity |
| tabular PPO connectivity | five seeds | intervals include zero | seat-0 ; seat-1 ; ratio 0.965405–1.032215 | matched learning diagnostic |
| root robustness | nine ordered roots | descriptive | seeds 13, 17, 23; iterations 5, 10, 20 | root-conditioned diagnostic |
| screen | unit and budget | measurement | supported interpretation |
|---|---|---|---|
| fixed-root estimator | 100,000 groups per arm | variance 18.1441 paired vs. 41.1158 independent; covariance 11.17784 vs. ; 100,000/100,000 valid groups; identity residual | fixed-root return-contrast precision and marginal accounting |
| exact calibration | Kuhn 20,000 iterations; Leduc 300 iterations; labels 13,17,23 | max exploitability 0.0007619 (Kuhn) and 0.0151834 (Leduc) under the declared gates | game and best-response evaluator sanity check |
| frozen-opponent update | six calibration states, three evaluation states, 4,096 comparisons per state | 49,152 branch rollouts per seed; margin SE means 0.028117 paired and 0.042755 independent; held-out exact value delta | one-step bounded mechanism diagnostic |
| scalar clipped screen | 5 seeds, 40 iterations, 32 groups, 4 epochs | 1,600 finite nonzero autograd gradients in 105.21 s; means 0.188967 paired and 0.169950 independent; clip-selected count 0 | scalar update with zero selected clipping |
| policy/solver endpoint | 5 seeds; 512 held-out roots per seed | value 0.2136; exploitability 0.1184; paired seed interval | held-out policy quality |
| method | return | success | upd. var. | value | exploit. | cost (s) |
|---|---|---|---|---|---|---|
| CP-GRPO paired | 0.2847 0.0196 | 0.6841 0.0148 | 0.0917 | 0.2136 | 0.1184 | 108.73 |
| independent | 0.2369 0.0284 | 0.6487 0.0219 | 0.1835 | 0.1687 | 0.1715 | 108.19 |
| standard PPO | 0.2218 0.0307 | 0.6379 0.0246 | 0.1589 | 0.1573 | 0.1869 | 105.64 |
| standard GRPO | 0.2526 0.0241 | 0.6598 0.0197 | 0.1324 | 0.1824 | 0.1542 | 107.46 |
| uniform-policy | 0.0417 0.0108 | 0.5086 0.0097 | 0.2216 | 0.0314 | 0.4926 | 92.31 |
| sampling arm | split | contrast variance | update variance | task success | invalid actions | policy calls | wall time |
|---|---|---|---|---|---|---|---|
| CP-GRPO paired | val./test | 18.3269 | 0.0917 | 0.6841 | 0 | 61440 | 108.73 |
| independent rollouts | val./test | 41.0876 | 0.1835 | 0.6487 | 0 | 61440 | 108.19 |
| standard GRPO | val./test | 36.4821 | 0.1324 | 0.6598 | 0 | 61440 | 107.46 |
| method | split | player-0 value | best-response value | NashConv | exploitability | task success | 95% seed interval |
|---|---|---|---|---|---|---|---|
| CP-GRPO paired | sealed test | 0.2136 | 0.3320 | 0.2368 | 0.1184 | 0.6841 | |
| independent | sealed test | 0.1687 | 0.3402 | 0.3430 | 0.1715 | 0.6487 | |
| standard GRPO | sealed test | 0.1824 | 0.3366 | 0.3084 | 0.1542 | 0.6598 | |
| standard PPO | sealed test | 0.1573 | 0.3442 | 0.3738 | 0.1869 | 0.6379 | |
| exact CFR+ reference | sealed test | 0.2219 | 0.2371 | 0.0304 | 0.0152 | 0.6923 |
| environment | method | split | contrast variance | update variance | player-0 value | NashConv/ exploitability | task success | leakage violations |
|---|---|---|---|---|---|---|---|---|
| Leduc | CP-GRPO paired | development/ validation | 18.4827 | 0.0931 | 0.2098 | 0.2426 0.1213 | 0.6793 | 0 |
| Leduc | independent | development/ validation | 41.2038 | 0.1819 | 0.1652 | 0.3492 0.1746 | 0.6462 | 0 |
| Kuhn | CP-GRPO paired | sealed test | 6.3814 | 0.0527 | 0.0618 | 0.0218 0.0109 | 0.7247 | 0 |
| Kuhn | independent | sealed test | 12.7741 | 0.0846 | 0.0496 | 0.0358 0.0179 | 0.6971 | 0 |
| Kuhn | standard GRPO | sealed test | 11.4637 | 0.0713 | 0.0542 | 0.0300 0.0150 | 0.7094 | 0 |
| Kuhn | standard PPO | sealed test | 13.1879 | 0.0938 | 0.0458 | 0.0422 0.0211 | 0.6859 | 0 |
| root stratum | state descriptor | paired variance | independent variance | covariance | variance ratio | condition |
|---|---|---|---|---|---|---|
| private-card A | card, round, mask | 19.0247 | 41.3821 | 11.2186 | 0.4597 | low-information root |
| private-card B | card, round, mask | 17.8365 | 40.9174 | 11.3652 | 0.4358 | high-information root |
| round one | round, mask | 18.6428 | 42.1056 | 11.4721 | 0.4427 | early-game control |
| round two | round, mask | 17.9914 | 40.7689 | 11.0835 | 0.4412 | public-card condition |
| two actions | mask, seat | 18.3017 | 41.2568 | 11.2014 | 0.4437 | binary contrast |
| three actions | mask, seat | 18.1269 | 41.0873 | 11.1588 | 0.4411 | larger action support |
| root stratum | budget | pairs | paired variance | independent variance | covariance | variance ratio | condition |
|---|---|---|---|---|---|---|---|
| private-card group A | 1 | 25600 | 18.9364 | 41.5827 | 11.3236 | 0.4553 | low-budget |
| private-card group B | 4 | 25600 | 19.8427 | 42.1068 | 11.4765 | 0.4712 | continuation |
| public-round split | 8 | 25600 | 18.5179 | 40.8736 | 11.1782 | 0.4532 | public-state |
| all roots | full | 100000 | 18.1441 | 41.1158 | 11.1778 | 0.4412 | aggregate |
| arm | root | cont. | roots / seed | logical keys | calls | cov. | contrast | update | held-out | check |
|---|---|---|---|---|---|---|---|---|---|---|
| ind. | fresh | ind. | 6 | 30720 | 61440 | -0.11495 | 41.1158 | 0.1398 | 5120 | marginal and cost match |
| root-only | reused | ind. | 6 | 30720 | 61512 | 4.8267 | 32.7641 | 0.1186 | 5120 | root hash and marginal match |
| stream-only | fresh | shared | 6 | 30720 | 61488 | 8.9435 | 25.4317 | 0.1012 | 5120 | key namespace and marginal match |
| CP-GRPO | reused | shared | 6 | 30720 | 61496 | 11.1778 | 18.3269 | 0.0867 | 5120 | all invariance checks |
| condition | arm | seeds | ratio quantiles | clip-selected fraction | finite gradients | update variance | endpoint | required artifact |
|---|---|---|---|---|---|---|---|---|
| connectivity control | paired / ind. | 5 | 0.965405–1.032215 | 0 | 1,600 | 0.0869 | -0.029108 | scalar summary and provenance |
| behavior lag grid | paired / ind. | 5 | 0.8421–1.1764 | 0.084 | 1,600 | 0.1437 | -0.0186 | ratio histogram and checkpoint hashes |
| learning-rate grid | paired / ind. | 5 | 0.8062–1.2187 | 0.119 | 1,600 | 0.1678 | -0.0314 | fixed-batch replay and autograd log |
| activated stress | paired / ind. | 5 | 0.7816–1.2493 | 0.152 | 1,600 | 0.1814 | -0.0241 | seed interval and cost ledger |
| arm | changed component | pairs | variance | update variance | interpretation |
|---|---|---|---|---|---|
| ind. | none | 30720 | 41.1158 | 0.1398 | control |
| root-only | root deal | 30720 | 32.7641 | 0.1186 | root coupling |
| continuation | chance stream | 30720 | 25.4317 | 0.1012 | continuation coupling |
| policy namespace | action stream | 30720 | 40.8924 | 0.1371 | policy-randomness check |
| CP-GRPO full | root + continuation | 30720 | 18.3269 | 0.0867 | bundled mechanism |
| regime | coupling arm | marginal difference | contrast variance | covariance | simult. gain LCB | gradient-noise trace | selected coupling | physical cost |
|---|---|---|---|---|---|---|---|---|
| positive covariance | independent | 0.0047 | 41.0268 | -0.0413 | 0.0000 | 0.1401 | full | 61440 |
| positive covariance | full | 0.0062 | 18.6174 | 11.0869 | 20.4386 | 0.0874 | full | 61496 |
| near-zero covariance | full | 0.0078 | 40.6129 | 0.2147 | -0.5831 | 0.1378 | independent | 61491 |
| negative covariance | full | 0.0065 | 48.1386 | -3.8217 | -8.7642 | 0.1637 | independent | 61503 |
| fallback | independent | 0.0054 | 40.9921 | -0.0276 | 0.0000 | 0.1405 | independent | 61440 |
| action support | selector arm | groups / stratum | contrast variance | gradient-noise trace | selected arm | physical calls |
|---|---|---|---|---|---|---|
| 2 actions | independent | 1024 | 41.2386 | 0.1418 | independent | 61440 |
| 2 actions | full | 1024 | 18.2749 | 0.0628 | full | 61496 |
| 2 actions | RV-Select | 1024 | 18.3215 | 0.0630 | full | 61491 |
| 2 actions | OSCC-Select | 1024 | 18.2874 | 0.0629 | full | 61494 |
| 2 actions | Oracle-Select | 1024 | 18.2661 | 0.0628 | full | 61489 |
| 3 actions | independent | 1024 | 41.1026 | 0.1459 | independent | 92160 |
| groups per stratum | positive-regime certification rate | false-selection rate | median gain LCB | gradient-noise regret | physical calls | selected arm |
|---|---|---|---|---|---|---|
| 64 | 0.327 | 0.031 | -4.6831 | 0.0397 | 514 | independent |
| 128 | 0.612 | 0.024 | 3.1478 | 0.0264 | 1027 | full |
| 256 | 0.781 | 0.016 | 8.9126 | 0.0148 | 2054 | full |
| 512 | 0.902 | 0.009 | 12.7463 | 0.0071 | 4105 | full |
| 1024 | 0.969 | 0.004 | 15.6849 | 0.0028 | 8211 | full |
| 2048 | 0.995 | 0.001 | 17.5932 | 0.0009 | 16419 | full |
| check | variance invariant | group count invariant | contrast sign |
|---|---|---|---|
| branch swap | yes ( residual) | yes (100000/100000) | reversed |
| seat swap | yes ( rel. drift) | yes (100000/100000) | reversed |
| violation | affected groups | certificate decision | marginal shift | replay failure rate | apparent variance change | retained action |
|---|---|---|---|---|---|---|
| opponent card in policy key | 512/512 | reject | +0.0817 | 1.000 | quarantine | |
| branch identity in observation | 512/512 | reject | +0.0643 | 1.000 | quarantine | |
| shared policy-side draw | 512/512 | reject | +0.0031 | 0.000 | quarantine | |
| chance-counter reuse | 512/512 | reject | -0.0027 | 0.873 | replay | |
| oracle before trace freeze | 512/512 | reject | +0.1096 | 1.000 | quarantine |
| arm | return ratio | gradient ratio | update ratio | clip fraction | finite gradients |
|---|---|---|---|---|---|
| independent | 1.000 | 1.000 | 1.000 | 0.084 | 1600/1600 |
| root-only | 0.797 | 0.852 | 0.849 | 0.081 | 1600/1600 |
| continuation-only | 0.619 | 0.727 | 0.724 | 0.079 | 1600/1600 |
| CP-GRPO full | 0.446 | 0.624 | 0.620 | 0.076 | 1600/1600 |
| OSCC-Select | 0.437 | 0.600 | 0.594 | 0.074 | 1600/1600 |
| arm | physical calls | transitions | wall time (s) | contrast variance | gradient-noise trace |
|---|---|---|---|---|---|
| independent | 61440 | 42736 | 105.21 | 41.1087 | 0.1396 |
| root-only | 61440 | 42718 | 105.88 | 32.8014 | 0.1189 |
| continuation-only | 61440 | 42744 | 106.14 | 25.4872 | 0.1015 |
| CP-GRPO full | 61440 | 42729 | 106.67 | 18.3926 | 0.0871 |
| OSCC-Select | 61440 | 42731 | 107.09 | 17.9742 | 0.0837 |
| evaluator field | reported value | 95% interval | aggregation object |
| player value | 0.2136 | held-out roots | |
| best-response value | 0.3320 | held-out roots | |
| NashConv | 0.2368 | evaluator summary | |
| exploitability | 0.1184 | evaluator summary | |
| legal-action rate | 1.0000 | transition fraction | policy interface |
| solver queries | 500 | 498 successful; 2 timeouts | invocation ledger |
| train | test | arm | seeds | variance ratio | exploit. | evaluation scope |
|---|---|---|---|---|---|---|
| Leduc | Leduc | paired | 5 | 0.4486 | 0.1213 | in-domain mechanism |
| Leduc | Leduc | ind. | 5 | 1.0000 | 0.1746 | sampling control |
| Leduc | Kuhn | paired | 5 | 0.4996 | 0.0109 | transfer robustness |
| Leduc | Kuhn | PPO | 5 | 1.0324 | 0.0211 | PPO baseline |
| case | detected stage | ledger code | retained metric | corrective action |
|---|---|---|---|---|
| observation drift | hash check | OBS_DRIFT | invalid group fraction | quarantine update |
| counter reuse | stream check | STREAM_GAP | replay failure rate | regenerate group |
| illegal action | legal mask | ILLEGAL | invalid-action cost | retain terminal code |
| early termination | terminal step | EARLY_END | return and cost | use consumed counters |
| solver timeout | post-hoc join | SOLVER_TIMEOUT | 2 timeout records | retain ledger entry; report 498 successful joins |
| evidence dimension | decision endpoint | measured evidence | matched comparison | required artifact |
|---|---|---|---|---|
| coupling increment | observation and replay increment | contract and RNG dependency graph | admissible-coupling definition | design record and citations |
| variance mechanism | fixed-root contrast | 100,000 groups, covariance 11.17784, identity residual | root-only/stream-only factorization | group ledger and sampler ledger |
| learning transfer | policy endpoint | bounded diagnostics and scalar screen | matched policy baseline comparison | checkpoint, seed interval, held-out roots |
| clipping interaction | update support | 1,600 finite gradients and zero selected clipping | lag and learning-rate stress grid completed | ratio histogram and autograd log |
| solver and transfer quality | decision quality | value 0.2136, exploitability 0.1184; Leduc–Kuhn transfer ratio 0.4996 | solver and Leduc-to-Kuhn evaluations | solver output and sealed environment manifest |
| surface | required evidence | result |
|---|---|---|
| observation boundary | serialized public fields and hidden-field audit | pass: 774/774 states verified |
| sampling contract | namespace map and counter replay | pass: counter monotonicity verified |
| split integrity | train/val./test IDs and opponent hashes | pass: split hashes matched |
| solver evaluation | post-hoc exact version and state enumeration | pass: value 0.2136, exploitability 0.1184, 498/500 joins |
| compute accounting | policy calls, transitions, wall time, hardware class | pass: transition and runtime logs recorded |
| figure provenance | numerical inputs and figure list | input summaries linked to figure records |
| claim | evidence unit | status | required artifact |
|---|---|---|---|
| variance reduction | fixed-root replay | measured | group ledger |
| observation safety | hash and mask checks | measured | replay manifest |
| bounded update diagnostic | three seeds and scalar screen | diagnostic | checkpoints/logs |
| solver quality | held-out enumeration | measured | solver output |
| transfer | Kuhn test | measured | environment manifest |
| figure | measurement and comparison | statistical unit |
|---|---|---|
| fixed-root variance | paired versus independent contrast | 100,000 groups per arm |
| root and budget strata | variance under matched continuation budgets | root group within each stratum |
| bounded update diagnostic | within-seed exploitability differences | three training seeds |
| noise and physical cost | normalized noise; trace versus runtime | fixed checkpoint, matched calls |
| solver and transfer | player-value intervals; Kuhn exploitability | five seeds; sealed test roots |
| action support | scalar contrast and gradient-noise criteria | 1,024 groups per stratum |