cs.LGSep 29, 2026

Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Authors: Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang

Organizations: KAIST · AITRICS

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

    Jun 23, 2026Wenyang Hu, Junxiang Jia, Zhen Shu +3Verifiable RewardsTrajectory-Level Credit

  2. Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

    May 8, 2026Tao Wang, Shuo Li, Yan Sun +2Reinforcement Learning With Verifiable RewardRepair-Based Group-Relative Policy Optimization

  3. Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

    Aug 4, 2026Yongshi Ye, Liang Zhang, Yidong Chen +2Reinforcement Learning With Verifiable RewardRepair-Based Group-Relative Policy Optimization