cs.AISep 27, 2026

Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification

Authors: Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Rui Chen, Daren Zha, Jun Xiao

Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China

Abstract

Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk ρ=0.08ρ=0.08, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost 1.16×1.16\times. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval [−0.0364,−0.0157][-0.0364,-0.0157]. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.

Explore similar work

CardsList
  1. Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

    Sep 28, 2026Christian Moya, Elliott Thornley, Guang LinReinforcement Learning With Verifiable RewardVerifiable Rewards

  2. Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

    Aug 11, 2026Jiamiao Liu, Dewen Qiao, Yu Zhang +1CertificationCovariate Shift

  3. A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

    Jun 7, 2026Xiaoli Yu, Jiamiao LiuConformal Risk ControlFinite-Sample Certificates