We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.
Figures & tables
Figure 1: MiniGrid tasks.
Figure 2: PPO results on MiniGrid . (a–b) Mean test reward over PPO rounds across five seeds; shaded regions show ±1 SEM. (c) At the final checkpoint, the lowest and highest 20 returns from 1,000 evaluation episodes per seed are pooled across five seeds as the Worst and Best subsets.
Clean
Mismatch-20%
Method
Acc. ↑
Δ Acc. ↑
Partial ↑
Acc. ↑
Δ Acc. ↑
Partial ↑
Pre-trained
0.4723 ± 0.0000
0.0000
0.4996 ± 0.0000
0.4723 ± 0.0000
0.0000
0.4996 ± 0.0000
Vanilla GRPO
0.4828 ± 0.0077
0.0105
0.5114 ± 0.0089
0.4731 ± 0.0094
0.0008
0.5001 ± 0.0089
Random 5%
0.5207 ± 0.0341
0.0484
0.5442 ± 0.0307
0.4103 ± 0.2026
-0.0620
0.4255 ± 0.2086
Random 10%
0.4511 ± 0.1175
-0.0212
0.4760 ± 0.1127
0.5133 ± 0.0431
0.0409
0.5318 ± 0.0455
Reward 5%
0.4516 ± 0.2104
-0.0208
0.4729 ± 0.2181
0.4443 ± 0.1691
-0.0281
0.4628 ± 0.1733
Table 1: GRPO results on GSM8K under clean and mismatch settings. Δ Acc. is measured relative to the shared pre-trained model. Results are reported as mean ± std. over five seeds, with the best result shown in bold. Given the consistent margins, we omit additional pairwise t -tests.
Figure 3: Empirical analysis of the self- and cross-term effect on GSM8K under Mismatch-20% . We define conflicted units as training units for which DTV and DTV-Loo make different keep/drop decisions; R denotes the fraction of units retained for visualization after trimming extreme score ranges, and IQR denotes the interquartile range (25th–75th percentiles). (a) Decomposition of the DTV score into the self- and cross-terms. (b) Filtering regions of DTV and DTV-Loo. (c) Self- and cross-term scores of conflicted units. (d) Cross-term contribution ∣cj∣/(sj+∣cj∣) for all and conflicted units. (e) Cross-term magnitude of conflicted units. (f) Drop ratios of DTV and the additional filtering induced by DTV-Loo.
Pass@1
Pass@16
Maj@16
Extractable ↑
Acc. given extractable ↑
Method
Score ↑
Δ↑
Score ↑
Δ↑
Score ↑
Δ↑
Pre-trained
0.1063
0.0000
0.3333
0.0000
0.2333
0.0000
0.7292
0.1457
Vanilla GRPO
0.1063
0.0000
0.4000
0.0667
0.2333
0.0000
0.7375
0.1441
DTV (Ours)
0.0917
-0.0146
0.3333
0.0000
0.2667
0.0334
0.7188
0.1275
DTV-Loo (Ours)
0.1333
0.0270
0.4333
0.1000
0.3667
0.1334
0.7521
0.1773
Table 2: GRPO results on AIME 2024 with a 32K maximum generation length. Δ denotes the change relative to the shared pre-trained model. Results are from a single training run for each method, with the best result shown in bold.
Clean
Mismatch-20%
Mismatch-40%
Method
AUC ↑
Acc. ↑
AUC ↑
Acc. ↑
AUC ↑
Acc. ↑
Vanilla DPO
0.5611 ± 0.0035
0.5388 ± 0.0073
0.5377 ± 0.0029
0.5014 ± 0.0044
0.4932 ± 0.0105
0.4301 ± 0.0160
Random 5%
0.5585 ± 0.0022
0.5449 ± 0.0023
0.5371 ± 0.0032
0.5013 ± 0.0152
0.4907 ± 0.0105
0.4234 ± 0.0143
Random 10%
0.5578 ± 0.0017
0.5421 ± 0.0063
0.5363 ± 0.0040
0.5015 ± 0.0100
0.4900 ± 0.0113
0.4332 ± 0.0122
Reward 5%
0.5620 ± 0.0016
0.5559 ± 0.0062
0.5465 ± 0.0028
0.5350 ± 0.0092
0.5083 ± 0.0089
0.4683 ± 0.0191
Reward 10%
0.5652 ± 0.0013
0.5647 ± 0.0063
0.5490 ± 0.0006
0.5443 ± 0.0048
0.5132 ± 0.0090
0.4857 ± 0.0147
Table 3: DPO preference-accuracy results on UltraFeedback under clean and mismatch settings. Results are reported as mean ± std. over five seeds, with the best result in each column shown in bold. † The final row reports two-sided paired t -test p -values comparing DTV-Loo with Reward 10% across matched seeds; significant differences ( p<0.05 ) are shown in bold.
Figure 4: DPO validation dynamics and training efficiency on UltraFeedback . The top row shows the change in validation accuracy relative to Vanilla DPO throughout training under Clean , Mismatch-20% , and Mismatch-40% conditions, averaged over five seeds; shaded regions indicate variability across seeds, and the horizontal dashed line denotes Vanilla DPO. The bottom row reports, for seed 0, the wall-clock time to reach the T95 target defined from the five-seed best validation performance; × NR denotes that the target is not reached within the training budget. DTV-Loo exhibits increasingly large and more persistent improvements as the mismatch rate increases, while also reaching the convergence target substantially faster under the mismatch settings.
Figure 5: Empirical analysis of the self-protection effect on UltraFeedback under Mismatch-20% . (a) Decomposition of the DTV Score into the self- and cross-terms. (b) Cross-term contribution ∣cj∣/(sj+∣cj∣) for all and conflicted units. (c) Visualization of self- and cross-term scores across samples in a representative batch (ranked by self-term). A single dominant sample with an extreme self-term score overwhelms the cross-term interactions across the entire batch, illustrating the self-protection mechanism in standard DTV.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Auxiliary signal
Fixed dataset
Validation set
Selection budget
Extra rollouts
Online valuation
Applied scenario
IIF ( 2025 )
Yes
No
No
Yes
No
Yes
PPO
LearnAlign ( 2026 )
Yes
Yes
No
Yes
Yes
No
GRPO / RLVR
GradAlign ( 2026 )
Yes
No
Yes
Yes
Yes
Yes
GRPO
DTV/DTV-Loo (Ours)
No
No
No
No
No
Yes
General
Appendix
Table 4: Comparison of additional requirements and applicability of gradient-based data valuation methods in reinforcement learning. The binary columns indicate whether each method uses an auxiliary valuation signal, assumes a fixed dataset, requires a validation set, uses an explicit selection budget, requires extra rollouts for valuation, or performs valuation online. The final column summarizes each method’s generalizability across reinforcement learning settings.
Hyperparameter
Empty-8x8
DoorKey-8x8
Parallel environments
16
16
Environment-step budget
160,000
1,000,000
Rollout steps per environment
128
640
Transitions per rollout round
2,048
10,240
Optimizer mini-batch size
64
256
PPO epochs per round
10
4
Appendix
Table 5: PPO training and evaluation configuration. The environment-step budget is the configured stopping target; collection ends after completing the final rollout round.
Hyperparameter
GSM8K
AIME
Prompts per update
4
128
Completions per prompt
4
8
Train micro-batch size
–
2
Maximum prompt / response length
256 / 768
1,024 / 8,192
Optimizer / peak learning rate
AdamW / 1×10−6
AdamW / 1×10−6
Learning-rate schedule
69-step warmup, cosine decay
cosine decay, no warmup
Appendix
Table 6: GRPO training and evaluation configurations on GSM8K and AIME 2024
Hyperparameter
SFT
DPO
Fine-tuning mode
Full parameter
LoRA ( r=64 , scaling =64 )
Train / evaluation batch size
2 / 2
8 / 8
Gradient accumulation steps
4
32
Maximum length
Target: 768
Prompt / response: 512 / 512
Optimizer / peak learning rate
AdamW / 1×10−5
AdamW / 1×10−6
Warmup / decay steps
100 / 1,500
10 / 115
Appendix
Table 7: SFT and DPO training configurations
Condition
Method
LiveBench-IF ↑
RB2 Precise IF ↑
IFBench P-Strict ↑
Clean
Vanilla DPO
0.2418 ± 0.0035
0.4938 ± 0.0285
0.1280 ± 0.0045
Random 5%
0.2413 ± 0.0046
0.4763 ± 0.0294
0.1227 ± 0.0090
Random 10%
0.2436 ± 0.0049
0.4813 ± 0.0442
0.1307 ± 0.0068
Reward 5%
0.2388 ± 0.0064
0.4800 ± 0.0203
0.1253 ± 0.0016
Reward 10%
0.2370 ± 0.0055
0.4613 ± 0.0191
0.1247 ± 0.0050
DTV (Ours)
0.2430 ± 0.0034
0.4938 ± 0.0198
0.1247 ± 0.0045
Appendix
Table 8: Clean downstream instruction-following performance of DPO methods. Results are mean ± sample standard deviation over five training seeds. These metrics are not used for training, filtering, or checkpoint selection; the best result in each condition and column is shown in bold.
Figure 6: Absolute preference-accuracy trajectories on UltraFeedback under Clean , Mismatch-20% , and Mismatch-40% training conditions. Curves show the five-seed mean and shaded regions show ±1 sample standard deviation.
Figure 7: Fixed-filtering-rate ablation under Mismatch-40% . DTV and DTV-Loo are compared with Random and Reward-based Filtering at 5%, 10%, 20%, and 40%. Curves show the five-seed mean, with shading denoting ±1 sample standard deviation.
Outcome-driven reinforcement learning offers a scalable way to post-train vision-language-action (VLA) policies from sparse task-success feedback. In common GRPO-based VLA post-training, one rollout-level advantage is applied to every action in the trajectory. A rollout that completes several valid stages but fails later can therefore penalize the actions that produced its earlier progress. We call this trajectory-level credit aliasing. Temporal GRPO addresses this problem by constructing detectable task stages, aligning each rollout with stage-specific action intervals, and comparing only rollouts that have entered the same stage. The resulting stage advantages are applied to their corresponding intervals in a single policy update. On RoboTwin 2.0, Temporal GRPO improves task success and sample efficiency, with consistent gains across task horizons. Controlled updates on LIBERO-Long preserve shared prerequisite stages and concentrate improvement at the first stage where rollout outcomes diverge.
Yao Zhou, Hang Gao, Fengge Wu +2
Institute of Software Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. Existing task-valuation methods mostly rely on snapshot-based signals such as current pass rate or reward, which estimate how solvable a task is under the current policy. However, tasks with similar current solvability can still differ substantially in how positively they respond to further training. We study this residual axis as task learnability: a regime-conditional measure of expected positive response to continued training under a fixed RL post-training regime. By analyzing per-task reward trajectories, we find that learnability is reproducible across independently sampled training contexts and predictive of downstream utility. To make this signal practical before training begins, we propose TrajVal, a lightweight probe-based estimator that approximates per-task learnability from a short probe run and two endpoint evaluations. TrajVal can be used either as a standalone static prior for task sampling or as a multiplicative prior for existing online schedulers. Experiments on mathematical and logical reasoning benchmarks across multiple model scales show that TrajVal improves data efficiency over uniform sampling and provides complementary gains when combined with online scheduling methods.
Offline reinforcement learning (RL) offers a path to policy improvement from logged data alone, using historical returns or other measurable outcomes as world feedback. A key difficulty is improving observed behavior without extrapolating beyond what the offline data supports. We propose \emph{counterfactual transport flows}, a source-conditioned trajectory refinement framework for offline decision-making guided by world feedback. Given a low-feedback candidate trajectory, we construct local preference pairs from offline data by retrieving nearby trajectories in latent trajectory space with higher task-specific feedback, and use them as weak supervision for conservative refinement. The framework learns instance-specific refinement directions: at inference time, a refinement strength parameter controls how far the candidate trajectory is transported, enabling a trade-off between preserving the original behavior and applying stronger improvement. Experiments on D4RL benchmarks, including AntMaze and MuJoCo tasks, show that our method improves behavior from historical returns as world feedback, while providing interpretable trajectory-level refinement paths.
Lena Krieger, Xuan Zhao, Zhuo Cao +3
1IAS-8, Forschungszentrum J ulich, Germany · 2LMU Munich, Munich Center for Machine Learning (MCML), Germany · Department of Computer Science, Aarhus University, Denmark.