cs.LGSep 30, 2026

Role-Adaptive Policy Optimization for Offline Reinforcement Learning

Authors: Seonvin Cho, Soohyun Choi, Songnam Hong

Organizations: Department of Electronic Engineering, Hanyang University Seoul, Republic of Korea

Abstract

Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning

    Aug 3, 2026Botao Dong, Longyang Huang, Ning Pang +1Diffusion PoliciesOffline Reinforcement Learning

  2. Ratio-Variance Regularized Policy Optimization

    May 26, 2026Yu Luo, Shuo Han, Yihan Hu +5Frictive Policy OptimizationTrust Region

  3. Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning

    Apr 21, 2026Yuan Zhuang, Yuexin Bian, Sihong He +7Critic LearningOff-Policy Learning