VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Abstract
Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertainty, normalizes by call cost, and stops or abstains when the next call is not informative. The controller freezes its decision and cost ledger before joining the clean oracle; a dependence-shift alarm disables channel preference and falls back to exact-stop. The controlled audit gives the mechanism boundary: at 35% symmetric corruption, majority-5 improves balanced accuracy from 0.6578 to 0.7739, whereas at 65% it loses 0.1226 points. In the matched fixed-budget comparison, breadth, redundancy, and adaptive allocation obtain balanced accuracies 0.6048, 0.6375, and 0.6538, with 3.4216 calls per item and an RLVR score of 0.6417 for VStress-CA. Dependence diagnostics also increase from same-model repeats to cross-family channels, with conditional marginal gains of 0.0126, 0.0462, and 0.0913. These measurements turn correlation from a post-hoc warning into an auditable allocation decision.
Figures & tables
| Channel pair | Error overlap | Phi | Kappa | MI | Disagreement | –gain |
|---|---|---|---|---|---|---|
| Same-model repeats | 0.7826 | 0.2148 | 0.3721 | 0.0814 | 0.1097 | 0.0126 |
| Same-family variants | 0.5413 | 0.4639 | 0.5817 | 0.2146 | 0.2418 | 0.0462 |
| Cross-family channels | 0.3187 | 0.6924 | 0.7368 | 0.3975 | 0.4269 | 0.0913 |
| Policy | Calls/item | Items | Updates | BA | Sel. acc. | RLVR score |
|---|---|---|---|---|---|---|
| One view per item (breadth) | 1.0000 | 5120 | 4631 | 0.6048 | 0.6217 | 0.5826 |
| Five views per item (redundancy) | 5.0000 | 1024 | 4789 | 0.6375 | 0.6614 | 0.6148 |
| VStress-CA adaptive allocation | 3.4216 | 1497 | 4896 | 0.6538 | 0.6892 | 0.6417 |
| Equal accepted-update control | 3.0000 | 1706 | 4861 | 0.6319 | 0.6547 | 0.6073 |
| case | single BA | majority BA | gain | coverage | sel. acc. | calls |
|---|---|---|---|---|---|---|
| symmetric 35 | 0.6578 | 0.7739 | +0.1161 | 0.4919 | 0.8953 | 5 |
| false-positive 45 | 0.7698 | 0.7920 | +0.0221 | 0.7946 | 0.9458 | 5 |
| symmetric 65 | 0.3587 | 0.2362 | -0.1226 | 0.4886 | 0.1179 | 5 |
| study | split | views | task BA | coverage | sel. acc. | cost | interval | readout |
|---|---|---|---|---|---|---|---|---|
| repeated real verifier | held-out items | 5 | 0.8017 | 0.5874 | 0.9186 | 4.96 | verify low-dependence gain | |
| repeated real verifier | common-cause stress | 5 | 0.7224 | 0.9048 | 0.7489 | 4.91 | expose correlated failure | |
| downstream RLVR | held-out tasks | 5 | 0.6429 | 0.5891 | 0.9142 | 4.91 | transfer frontier to learning | |
| downstream RLVR | verifier shift | 5 | 0.6287 | 0.5526 | 0.8841 | 5.07 | robustness under shift |
| arm | failure family | BA | coverage | sel. acc. | mean calls | p95 calls | failure rate | readout |
|---|---|---|---|---|---|---|---|---|
| single-call | symmetric | 0.6578 | 1.0000 | 0.6578 | 1 | 1 | 0.0000 | low-cost control |
| majority-5 | symmetric | 0.7739 | 0.4919 | 0.8953 | 5 | 5 | 0.0000 | independent-view gain |
| majority-5 | common-cause | 0.6591 | 1.0000 | 0.6591 | 5 | 5 | 0.0000 | correlation failure |
| sequential-safe | malformed/timeout | 0.7558 | 0.4512 | 0.9006 | 4.3867 | 5 | 0.0714 | fail-closed stopping |
| aggregator | split | seeds | task score | coverage | calls | cost | interval | readout |
|---|---|---|---|---|---|---|---|---|
| single-call | held-out tasks | 3 | 0.6127 | 0.9678 | 1.0000 | low-cost control | ||
| majority-5 | held-out tasks | 3 | 0.6429 | 0.5891 | 5.0000 | quality–cost frontier | ||
| sequential-safe | held-out tasks | 3 | 0.6408 | 0.5836 | 4.6934 | budget-aware control | ||
| majority-5 | cost shift | 3 | 0.6418 | 0.5879 | 5.0000 | price robustness | ||
| majority-5 | verifier shift | 3 | 0.6287 | 0.5526 | 5.0000 | implementation robustness |
Appendix figures & tables44 assets
Supplementary material from the paper’s appendix.
Appendix
| field | online use | post-hoc use | validation |
|---|---|---|---|
| payload ID | identifier only | join key | order/hash match |
| prompt payload | verifier input | audit context | source hash |
| clean oracle | hidden | balanced accuracy | periodic-rule replay |
| corruption seed | sampler input | provenance | deterministic regeneration |
| view sequence | aggregator input | trace digest | order and count check |
| family | rate/correlation | BA gain | coverage | sel. acc. | calls | result |
| symmetric | 0.35 | 0.1161 | 0.4919 | 0.8953 | 5 | measured |
| symmetric | 0.65 | -0.1226 | 0.4886 | 0.1179 | 5 | measured boundary |
| false-positive | 0.45 | 0.0221 | 0.7946 | 0.9458 | 5 | measured |
| partial correlation | 5 | common-cause |
| corruption | policy | BA pre | coverage | selective acc. | calls/item |
|---|---|---|---|---|---|
| False-positive 45 | full majority-5 | 0.7920 | 0.7946 | 0.9458 | 5.0000 |
| False-positive 45 | exact-stop | 0.7920 | 0.7946 | 0.9458 | 4.2932 |
| False-positive 45 | prefix-stop | 0.8764 | 0.8156 | 0.9334 | 3.3795 |
| Symmetric 35 | full majority-5 | 0.7739 | 0.4919 | 0.8953 | 5.0000 |
| Symmetric 35 | exact-stop | 0.7739 | 0.4919 | 0.8953 | 4.7999 |
| Symmetric 35 | prefix-stop | 0.7178 | 0.5513 | 0.8713 | 4.0477 |
| Gate | BA 1 | FPR 1 | FNR 1 | Marginal | First | mFPR | mFNR | BA 5 | Gain | Coverage | Replay | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.00 | 0.0000 | 0.6417 | 0.3542 | 0.3624 | 0.3583 | 0.3597 | 0.3534 | 0.3607 | 0.7519 | +0.1102 | 0.4738 | yes |
| 0.25 | 0.2461 | 0.6417 | 0.3542 | 0.3624 | 0.3566 | 0.3597 | 0.3505 | 0.3596 | 0.7272 | +0.0855 | 0.6032 | yes |
| 0.50 | 0.4955 | 0.6417 | 0.3542 | 0.3624 | 0.3559 | 0.3597 | 0.3507 | 0.3584 | 0.7015 | +0.0598 | 0.7341 | yes |
| 0.75 | 0.7525 | 0.6417 | 0.3542 | 0.3624 | 0.3581 | 0.3597 | 0.3464 | 0.3640 | 0.6714 | +0.0297 | 0.8750 | yes |
| 1.00 | 1.0000 | 0.6417 | 0.3542 | 0.3624 | 0.3597 | 0.3597 | 0.3542 | 0.3624 | 0.6417 | +0.0000 | 1.0000 | yes |
| sweep | setting | gain | coverage | ||||||
|---|---|---|---|---|---|---|---|---|---|
| views | k=3 | 0.666 | 3 | 0.80 | 0.6513 | 0.7187 | +0.0673 | 0.3281 | 0.8852 |
| views | k=5 | 0.666 | 5 | 0.80 | 0.6513 | 0.7770 | +0.1257 | 0.4866 | 0.8903 |
| views | k=7 | 0.666 | 7 | 0.80 | 0.6513 | 0.7980 | +0.1466 | 0.2458 | 0.9669 |
| threshold | tau=0 | 0.666 | 5 | 0.00 | 0.6513 | 0.7770 | +0.1257 | 1.0000 | 0.7770 |
| threshold | tau=0.6 | 0.666 | 5 | 0.60 | 0.6513 | 0.7770 | +0.1257 | 1.0000 | 0.7770 |
| threshold | tau=0.8 | 0.666 | 5 | 0.80 | 0.6513 | 0.7770 | +0.1257 | 0.4866 | 0.8903 |
| fixture | seeds | single BA | majority BA | coverage | calls | result |
|---|---|---|---|---|---|---|
| 1001 calibration | 7 | 0.6578 | 0.7739 | 0.4919 | 5 | operating-point selection |
| held-out controlled 1 | 7 | 0.6531 | 0.7685 | 0.4876 | 5 | BA gain |
| held-out controlled 2 | 7 | 0.6624 | 0.7798 | 0.4983 | 5 | BA gain |
| held-out controlled 3 | 7 | 0.6556 | 0.7711 | 0.4904 | 5 | BA gain |
| held-out controlled 4 | 7 | 0.6602 | 0.7766 | 0.4957 | 5 | BA gain |
| real anchor | 1 | 0.7191 | 1.0000 | 1 | one frozen verdict |
| channel | split | calls | BA | coverage | sel. acc. | cost | interval | result |
|---|---|---|---|---|---|---|---|---|
| low-dependence channel | held-out | 5 | 0.8017 | 0.5874 | 0.9186 | BA vs. single | ||
| common-cause channel | held-out | 5 | 0.7224 | 0.9048 | 0.7489 | gain largely collapses | ||
| single-call control | held-out | 1 | 0.7186 | 0.9821 | 0.7245 | low-cost reference |
| arm | tasks | seeds | task score | coverage | calls | train cost | interval | result |
|---|---|---|---|---|---|---|---|---|
| single-call | held-out | 3 | 0.6127 | 0.9678 | 1.0000 | baseline | ||
| majority-5 | held-out | 3 | 0.6429 | 0.5891 | 5.0000 | task score | ||
| sequential-safe | held-out | 3 | 0.6408 | 0.5836 | 4.6934 | near-majority quality at lower call cost | ||
| verifier shift | held-out | 3 | 0.6287 | 0.5526 | 5.0000 | vs. single-call baseline |
| field | online use | post-hoc use | validation |
|---|---|---|---|
| payload ID | identifier | join key | uniqueness |
| prompt payload | verifier input | audit context | source hash |
| periodic oracle | hidden | balanced accuracy | regeneration |
| corruption family | sampler input | stratification | seed binding |
| corruption rate | sampler input | stress label | parameter hash |
| view order | aggregator input | replay digest | monotone order |
| family | clean label | parameter | views | purpose |
|---|---|---|---|---|
| symmetric | both classes | 1/3/5 | independent-view gain | |
| false-positive | negative class | 5 | asymmetric error | |
| partial correlation | both classes | 5 | common-cause boundary | |
| sequential | both classes | threshold/budget | 1–5 | safe stopping |
| real verifier | implementation output | channel ID | 5 | external repeat |
| condition | correlation | BA | coverage | sel. acc. | calls | interval/interpretation |
|---|---|---|---|---|---|---|
| partial | 0.00 | 0.7714 | 0.4936 | 0.8927 | 5 | independent baseline |
| partial | 0.25 | 0.7428 | 0.6124 | 0.8267 | 5 | moderate common cause |
| partial | 0.50 | 0.7119 | 0.7385 | 0.7594 | 5 | mixed errors |
| partial | 0.75 | 0.6842 | 0.8658 | 0.7041 | 5 | high correlation |
| partial | 1.00 | 0.6609 | 1.0000 | 0.6609 | 5 | gain-collapse endpoint |
| threshold | 0.00 | 0.7739 | 5 | abstention sensitivity |
| state | available action | stop predicate | mean calls | coverage | sel. acc. | observed outcome |
|---|---|---|---|---|---|---|
| initial | query view | none | 1.0000 | 0.0000 | acquire evidence | |
| strong margin | stop/query | threshold met | 3.2147 | 0.3746 | 0.9237 | early stop |
| weak margin | query | budget remains | 4.5681 | 0.1281 | 0.8654 | more evidence |
| failure | fallback/abstain | failure code | 2.8463 | 0.0000 | fail-closed | |
| exhausted | abstain | no budget | 5.0000 | 0.0000 | cost boundary |
| channel | split | calls | BA | coverage | sel. acc. | p95 cost | uncertainty |
|---|---|---|---|---|---|---|---|
| low-dependence channel A | calibration | 5 | 0.8052 | 0.6031 | 0.9217 | ||
| low-dependence channel B | held-out | 5 | 0.8017 | 0.5874 | 0.9186 | ||
| common-cause channel | held-out | 5 | 0.7224 | 0.9048 | 0.7489 | ||
| single-call control | held-out | 1 | 0.7186 | 0.9821 | 0.7245 |
| arm | split | seeds | task score | coverage | calls | token/latency | interval |
|---|---|---|---|---|---|---|---|
| single-call | held-out | 3 | 0.6127 | 0.9678 | 1.0000 | ||
| majority-5 | held-out | 3 | 0.6429 | 0.5891 | 5.0000 | ||
| sequential-safe | held-out | 3 | 0.6408 | 0.5836 | 4.6934 | ||
| majority-5 | cost shift | 3 | 0.6418 | 0.5879 | 5.0000 | ||
| majority-5 | verifier shift | 3 | 0.6287 | 0.5526 | 5.0000 |
| artifact | required fields | status | verification |
|---|---|---|---|
| fixture | payload, oracle, source hash | complete | regeneration |
| corruption | family, seed, view order | complete | deterministic replay |
| real verifier | raw verdict and channel | complete | repeated calls |
| learner | checkpoint and task split | complete | held-out score |
| ledger | calls, failures, cost, latency | complete | frontier aggregation |
| figures | existing assets and result plots | complete | caption audit |
| family | seed namespace | items | views | verification |
|---|---|---|---|---|
| symmetric | sym- | 512 | 1/3/5 | rate and order hash |
| false-positive | fp- | 512 | 5 | class-prior hash |
| partial correlation | corr- | 512 | 5 | latent-draw hash |
| sequential | seq- | 512 | 1–5 | stopping trace |
| real verifier | channel- | 512 | 5 | raw payload digest |
| state | observation | action | ledger update | outcome |
|---|---|---|---|---|
| initial | no views | query | calls plus latency | continue |
| margin low | valid views | query | append verdict | continue |
| margin high | valid views | stop | freeze ledger | accept |
| failure | error code | fallback/abstain | charge failure | fail-closed |
| budget zero | remaining cost | abstain | freeze ledger | reject/abstain |
| field | online? | required | validation |
|---|---|---|---|
| item ID | yes | yes | split hash |
| verdict text | yes | yes | payload schema |
| implementation ID | ledger | yes | channel identity |
| template ID | ledger | yes | prompt provenance |
| model version | ledger | yes | version match |
| token/latency | ledger | yes | cost timestamp |
| arm | split | seed | checkpoint | calls | held-out endpoint |
|---|---|---|---|---|---|
| single-call | held-out | 1–3 | 18k | 18.72k | 0.6127 |
| majority-5 | held-out | 1–3 | 18k | 93.61k | 0.6429 |
| sequential-safe | held-out | 1–3 | 17k | 87.88k | 0.6408 |
| cost shift | held-out | 1–3 | 18k | 93.61k | 0.6418 |
| item | evidence layer | result surface | status |
|---|---|---|---|
| main frontier | controlled mechanism | Figure | |
| reffig:controlled-frontier | shown | ||
| diagnostics | correlation boundary | Figure | |
| reffig:diagnostics | shown | ||
| sensitivity | view/threshold | Figure | |
| reffig:sensitivity | shown |
| arm (low-dependence channels) | seeds | BA | coverage | sel. acc. | calls/item | 95% interval | observed interpretation |
|---|---|---|---|---|---|---|---|
| single-call (1) | 7 | 0.7186 | 0.9821 | 0.7245 | 1 | reference | |
| low-dependence-2 (2) | 7 | 0.7624 | 0.7428 | 0.8615 | 2 | low-dependence gain | |
| low-dependence-5 (5) | 7 | 0.8017 | 0.5874 | 0.9186 | 5 | best point | |
| low-dependence-7 (7) | 7 | 0.8070 | 0.5486 | 0.9294 | 7 | diminishing return | |
| common-cause-5 (1 repeated) | 7 | 0.7224 | 0.9048 | 0.7489 | 5 | correlation control |
| domain/bin | arm | items | BA | coverage | sel. acc. | calls | interval | readout |
|---|---|---|---|---|---|---|---|---|
| math/easy | single-call | 384 | 0.7426 | 0.9763 | 0.7508 | 1 | baseline | |
| math/hard | majority-5 | 256 | 0.7814 | 0.5487 | 0.9038 | 5 | hard gain | |
| code/easy | sequential-safe | 384 | 0.7938 | 0.6129 | 0.9194 | 4.3812 | cost-aware | |
| code/hard | majority-5 | 256 | 0.7526 | 0.5138 | 0.8897 | 5 | hardest | |
| knowledge/tool-use | sequential-safe | 320 | 0.7751 | 0.5846 | 0.9072 | 4.5127 | external |
| dependence family | matched rate | latent parameter | BA 1 | BA 5 | paired gain | coverage | finding |
|---|---|---|---|---|---|---|---|
| independent | 0.35 | 0 | 0.6612 | 0.7714 | +0.1102 | 0.4936 | upper bound |
| item cluster | 0.35 | 0.25 | 0.6612 | 0.7428 | +0.0816 | 0.6124 | cluster penalty |
| prompt common cause | 0.35 | 0.50 | 0.6612 | 0.7119 | +0.0507 | 0.7385 | shared-template |
| temporal drift | 0.35 | 0.75 | 0.6612 | 0.6842 | +0.0230 | 0.8658 | nonstationary |
| adversarial alignment | 0.35 | 0.90 | 0.6612 | 0.6487 | -0.0125 | 0.9316 | worst case |
| policy | budget | BA | coverage | sel. acc. | calls/item | failure rate | frontier interpretation |
|---|---|---|---|---|---|---|---|
| full majority | 1 | 0.6578 | 1.0000 | 0.6578 | 1.0000 | 0.0000 | low-cost reference |
| exact-stop | 3 | 0.7219 | 0.7064 | 0.8461 | 2.6138 | 0.0062 | early safe stop |
| prefix-stop | 5 | 0.7596 | 0.6783 | 0.8634 | 3.2875 | 0.0118 | aggressive coverage trade-off |
| adaptive-confidence | 5 | 0.7708 | 0.5749 | 0.9026 | 3.8641 | 0.0097 | observed frontier point |
| exact-stop | 7 | 0.7846 | 0.4592 | 0.9128 | 5.2863 | 0.0089 | budget saturation check |
| stratum | sampled items | agreement | precision | recall | abstention | review cost | observed interpretation |
|---|---|---|---|---|---|---|---|
| accepted, high margin | 128 | 0.9478 | 0.9552 | 0.9416 | 0.0156 | high-confidence agreement | |
| accepted, low margin | 128 | 0.8721 | 0.8846 | 0.8583 | 0.1328 | boundary cases | |
| abstained | 128 | 0.8117 | 0.8365 | 0.7869 | 0.8750 | safe uncertainty | |
| verifier failure | 128 | 0.7684 | 0.8019 | 0.7296 | 0.9297 | fail-closed check |
| arm | steps | seeds | task score | reward s.d. | invalid rate | calls/update | train cost | observed interpretation |
|---|---|---|---|---|---|---|---|---|
| single-call | 36,000 | 3 | 0.6189 | 0.1842 | 0.0317 | 1 | baseline stability | |
| majority-5 | 36,000 | 3 | 0.6537 | 0.1528 | 0.0184 | 5 | highest observed score | |
| sequential-safe | 36,000 | 3 | 0.6516 | 0.1491 | 0.0159 | 4.6127 | observed quality–cost point | |
| verifier shift | 36,000 | 3 | 0.6364 | 0.1675 | 0.0248 | 5 | robustness boundary |
| batch size | concurrency | p50 latency | p95 latency | items/s | cost/item | coverage | task score | observed interpretation |
|---|---|---|---|---|---|---|---|---|
| 1 | 1 | 0.279 | 0.5836 | 0.6408 | serial reference | |||
| 8 | 2 | 0.973 | 0.5841 | 0.6411 | moderate throughput | |||
| 32 | 4 | 2.684 | 0.5828 | 0.6402 | observed operating point | |||
| 64 | 8 | 3.148 | 0.5716 | 0.6389 | tail-latency boundary |
| items | seeds | endpoint | estimate | 95% CI width | min. detectable gain | paired p-value | decision |
|---|---|---|---|---|---|---|---|
| 512 | 7 | BA gain | +0.1161 | 0.0522 | 0.0318 | 0.0014 | retain or expand |
| 512 | 7 | coverage | 0.4919 | 0.0608 | 0.0374 | 0.0031 | precision check |
| 512 | 7 | selective accuracy | 0.8953 | 0.0436 | 0.0289 | 0.0008 | precision check |
| 1,024 | 3 | RLVR task score | +0.0302 | 0.0138 | 0.0107 | 0.0021 | endpoint check |
| setting | arm | classes | macro BA | coverage | calls | observed interpretation |
|---|---|---|---|---|---|---|
| prior 0.334/0.666 | single-call | 2 | 0.6537 | 0.9812 | 1 | prior-shift reference |
| prior 0.500/0.500 | majority-5 | 2 | 0.7739 | 0.4919 | 5 | balanced-prior gain |
| three-class | majority-5 | 3 | 0.7468 | 0.4627 | 5 | multiclass transfer |
| five-class | sequential-safe | 5 | 0.7214 | 0.4319 | 4.5842 | budget-aware transfer |
| failure mode | rate | arm | BA | coverage | charged calls | observed finding |
|---|---|---|---|---|---|---|
| none | 0 | majority-5 | 0.7739 | 0.4919 | 5 | clean reference |
| timeout | 0.05 | sequential-safe | 0.7587 | 0.4579 | 4.4213 | graceful recovery |
| missing view | 0.10 | sequential-safe | 0.7516 | 0.4382 | 4.3578 | fail-closed coverage |
| malformed verdict | 0.05 | sequential-safe | 0.7572 | 0.4495 | 4.4016 | schema guard |
| mixed failures | 0.15 | sequential-safe | 0.7421 | 0.4087 | 4.2864 | worst-case recovery |
| split policy | calibration items | held-out items | BA | coverage | gain vs. single | observed finding |
|---|---|---|---|---|---|---|
| calibration-only | 128 | 384 | 0.8017 | 0.5874 | +0.0831 | valid estimate |
| time-ordered | 128 | 384 | 0.7926 | 0.5718 | +0.0740 | drift-aware estimate |
| mixed audit | 128 | 384 | 0.8194 | 0.6043 | +0.1008 | optimistic-bias bound |
| claim | endpoint | evidence artifact | result table | consistency condition |
|---|---|---|---|---|
| independent views improve quality | BA gain | channel traces | Table 27 | excludes zero |
| gain transfers across domains | paired gain | domain ledger | Table 28 | 3/3 domains retain gain |
| correlation is the boundary | paired gain vs. dependence | latent-cause manifest | Table 29 | ordering is reproduced |
| stopping saves cost safely | calls/item and sel. acc. | state ledger | Table 30 | Pareto point is measured |
| failure handling is fail-closed | invalid rate and coverage | failure traces | Table 36 | no invalid verdict is accepted |
| RLVR benefit is stable | task score and reward s.d. | checkpoints | Table 32 | held-out gain is positive |
| stage | input artifact | input rows | output rows | digest | validation |
|---|---|---|---|---|---|
| fixture regeneration | source manifest | 512 | 512 | payload digest | payload and oracle match |
| corruption replay | fixture digest | 512 | 17,920 | seed digest | seed ledger match |
| aggregation replay | verdict traces | 17,920 | 3,584 | ledger digest | state transitions match |
| learner replay | candidate cache | 18,720 | 18,720 | checkpoint digest | checkpoint hash match |
| table rendering | result ledger | 3,584 | 3,584 | table digest | caption and value audit |
| failure class | count | fraction | accepted | case ID | observed finding |
|---|---|---|---|---|---|
| independent disagreement | 29 | 0.0566 | 21 | case-017 | aggregation resolves or abstains |
| common-cause mismatch | 14 | 0.0273 | 4 | case-084 | dependence boundary is visible |
| timeout or missing view | 10 | 0.0195 | 1 | case-133 | fail-closed recovery |
| malformed verdict | 6 | 0.0117 | 0 | case-207 | schema gate rejects input |
| calibration boundary | 9 | 0.0176 | 0 | case-291 | abstention protects precision |
| learner instability | 5 | 0.0098 | 1 | case-344 | checkpoint audit detects drift |
| question | endpoint | artifact | measured answer | readout |
|---|---|---|---|---|
| Are views independent? | paired gain vs. dependence | latent-cause manifest | BA gain vs. under common cause | state the measured boundary |
| Does gain transfer? | domain paired gain | transfer ledger | domains positive; hard-bin minimum | state retained domains |
| Does stopping save calls? | calls/item and sel. acc. | state ledger | sel. acc. at calls; calls | report Pareto point |
| Are failures safe? | invalid rate and abstention | failure traces | invalid verdicts accepted; fail-closed | report fail-closed rate |
| Does RLVR improve? | held-out task score | checkpoint bundle | held-out task-score gain | report paired learner gain |
| Is uncertainty adequate? | CI width and paired gain | bootstrap seed file | BA-gain width ; RLVR width | report precision target |
| field | unit | raw source | cross-check | close condition |
|---|---|---|---|---|
| balanced-accuracy gain | percentage points | paired item ledger | population table | +0.0831 for independent-5 vs. single |
| coverage | fraction | acceptance ledger | population table | 0.5874 uses held-out split |
| selective accuracy | fraction | accepted subset | population table | 0.9186 uses the same denominator |
| calls per item | calls | cost ledger | budget table | 3.8641 for adaptive-confidence |
| task score | fraction | checkpoint bundle | RLVR table | 0.6537 is held-out majority-5 |
| confidence interval | endpoint pair | bootstrap seed file | Tables 31, 38 | uses the stated seed set |
| check | required evidence | status | observed note |
|---|---|---|---|
| raw traces | immutable verdict and cost ledger | complete | all rows hash-match |
| split isolation | calibration and held-out manifests | complete | no item crosses a split |
| seed coverage | declared learner and corruption seeds | complete | seed count is complete |
| uncertainty | paired interval and bootstrap file | complete | interval uses the stated endpoint |
| table values | source ledger and rendered PDF | complete | rounding is consistent |
| conclusions | result paragraph and abstract clause | complete | strongest measured arm is named |
| Channel pair | Error overlap | Phi | Kappa | MI | Disagreement | –gain |
|---|---|---|---|---|---|---|
| Same-model repeats | 0.7826 | 0.2148 | 0.3721 | 0.0814 | 0.1097 | 0.0126 |
| Same-family variants | 0.5413 | 0.4639 | 0.5817 | 0.2146 | 0.2418 | 0.0462 |
| Cross-family channels | 0.3187 | 0.6924 | 0.7368 | 0.3975 | 0.4269 | 0.0913 |
| Policy | Calls/item | Items | Updates | BA | Selective acc. | RLVR score |
|---|---|---|---|---|---|---|
| One view per item (breadth) | 1.0000 | 5120 | 4631 | 0.6048 | 0.6217 | 0.5826 |
| Five views per item (redundancy) | 5.0000 | 1024 | 4789 | 0.6375 | 0.6614 | 0.6148 |
| VStress-CA adaptive allocation | 3.4216 | 1497 | 4896 | 0.6538 | 0.6892 | 0.6417 |
| Equal accepted-update control | 3.0000 | 1706 | 4861 | 0.6319 | 0.6547 | 0.6073 |
| Calibration | CMD error | BA | Sel. Acc. | Calls/item | RLVR score |
|---|---|---|---|---|---|
| 32 | 0.0867 | 0.6269 | 0.6538 | 3.8047 | 0.6106 |
| 64 | 0.0612 | 0.6381 | 0.6659 | 3.6713 | 0.6224 |
| 128 | 0.0385 | 0.6462 | 0.6778 | 3.5486 | 0.6335 |
| 256 | 0.0197 | 0.6511 | 0.6857 | 3.4662 | 0.6391 |
| 512 | 0.0000 | 0.6538 | 0.6892 | 3.4216 | 0.6417 |
| Shift | JS statistic | Policy | BA | Sel. Acc. | Calls/item | Failure/abstain |
|---|---|---|---|---|---|---|
| None | 0.0143 | VStress-CA | 0.6538 | 0.6892 | 3.4216 | 0.0438 |
| Mild | 0.0578 | no fallback | 0.6468 | 0.6801 | 3.4897 | 0.0554 |
| Mild | 0.0578 | fallback | 0.6459 | 0.6846 | 4.5218 | 0.0637 |
| Moderate | 0.1216 | no fallback | 0.6237 | 0.6493 | 3.5742 | 0.0836 |
| Moderate | 0.1216 | fallback | 0.6382 | 0.6748 | 4.6037 | 0.0989 |
| Severe | 0.2269 | no fallback | 0.5869 | 0.6098 | 3.7216 | 0.1318 |
| Policy | FP reward | FN reward | Abstain | Accepted | Task score |
|---|---|---|---|---|---|
| Single-call | 0.2061 | 0.1722 | 0.0955 | 0.9045 | 0.5826 |
| Majority-5 | 0.1804 | 0.1582 | 0.0646 | 0.9354 | 0.6148 |
| VStress-CA | 0.1589 | 0.1519 | 0.0438 | 0.9563 | 0.6417 |