cs.AISep 28, 2026

SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL

Authors: Zenghuang Fu, Ningqi Chen, Mingda Jia, Xiaofeng Han, Zhaoyang Li, Qiuyuan Ai, Zelong Zheng, Haoyu Wu, +5 more

Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · The University of Hong Kong · Peking University · Mininglamp Technology · Key Laboratory of Computing Power Network and Information Security, Ministry of Education; Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences) · Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science

Abstract

Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh siblings are sampled after selection. When that statistic is associated with return, branch values can reflect selection history as well as continuation quality, even for a shared parent. We propose Selective-Inference Policy Optimization (\SIPO{}), which incorporates this distinction into tree-based credit estimation. Its scale-free branch criterion keeps generation scores and sibling penalties on a consistent relative scale; exchangeable branching supplies multiple fresh continuations from each selected parent; and order-statistic correction adjusts retained incumbent values using selection rank and the estimated score--outcome association. These mechanisms preserve the leaf budget and the host policy optimisation objective. Across seven QA benchmarks using Qwen3-4B, Qwen3-8B, and Qwen2.5-7B, \SIPO{} achieves the highest reported multi-hop and single-hop averages among the compared methods. On Qwen3-8B, it improves these averages over AT\textsuperscript{2}PO by 1.311.31 and 1.071.07 percentage points, respectively, and ranks first on six of seven benchmarks. Component ablations evaluate the individual and combined changes, while early-training paired diagnostics show a selected--fresh value gap alongside a near-zero fresh--fresh reference. Together, these results support accounting for selection history when constructing and evaluating search-agent rollouts. Our code is available at https://github.com/Zenghuang-Fu/SIPO

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CATPO: Critique-Augmented Tree Policy Optimization

    Jun 6, 2026Ayush Singh, Umang Goyal, Ankur DahiyaReinforcement Learning With Verifiable RewardTrees

  2. Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

    Jul 15, 2026Bowei He, Yankai Chen, Xiaokun Zhang +1Frictive Policy OptimizationOffline Reinforcement Learning