Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite secondary-verification budget; its action trace is frozen before clean labels are joined. Simultaneous finite-sample bounds then certify selected harmful risk, coverage, and verifier-call rate over a predeclared policy family. Conditional Hoeffding-Azuma bounds account for the dependence induced by shared budgets, and rollout or verifier changes initiate a new certification stage. After admission, a bounded trust-clip-KL actuator controls magnitude. We evaluate two models on two reasoning benchmarks against static RLVR, matched-random selection, confidence thresholding, noise correction, and verifier augmentation. On Qwen3.5-0.8B and GSM8K at target risk ρ=0.08, RC-VAPO achieves 74.1% accuracy, selected harmful risk 0.0697, coverage 0.4125, and relative verifier cost 1.16×. At matched coverage and update magnitude, its selected-risk difference from matched random is -0.0260 with paired 95% interval [−0.0364,−0.0157]. Across asymmetric, confidence-dependent, and correlated-verifier noise, the certificate is satisfied on 57 of 60 independent runs. These comparisons isolate informative directional selection from proposal suppression, update shrinkage, and additional verifier computation.
School of Cyber Security and Information Law Chongqing University of Posts and Telecommunications Chongqing 400065, China · Department of Information, Xinqiao Hospital Army Medical University (Third Military Medical University) Chongqing 400037, China