We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated trajectories without explicit validation signals, making conventional influence-based methods inapplicable. We propose Dynamic Trajectory Valuation (DTV), a simple and efficient framework that estimates trajectory utility at the mini-batch level and filters detrimental trajectories based solely on gradient information. By operating at the optimization level, DTV integrates seamlessly with existing reinforcement learning pipelines with minimal overhead. Extensive experiments across diverse settings, including PPO, GRPO, and DPO, demonstrate that DTV consistently improves performance, enhances data efficiency, and stabilizes optimization.
Figures & tables
Figure 1: MiniGrid tasks.
Figure 2: PPO results on MiniGrid . (a–b) Mean test reward over PPO rounds across five seeds; shaded regions show ±1 SEM. (c) At the final checkpoint, the lowest and highest 20 returns from 1,000 evaluation episodes per seed are pooled across five seeds as the Worst and Best subsets.
Clean
Mismatch-20%
Method
Acc. ↑
Δ Acc. ↑
Partial ↑
Acc. ↑
Δ Acc. ↑
Partial ↑
Pre-trained
0.4723 ± 0.0000
0.0000
0.4996 ± 0.0000
0.4723 ± 0.0000
0.0000
0.4996 ± 0.0000
Vanilla GRPO
0.4828 ± 0.0077
0.0105
0.5114 ± 0.0089
0.4731 ± 0.0094
0.0008
0.5001 ± 0.0089
Random 5%
0.5207 ± 0.0341
0.0484
0.5442 ± 0.0307
0.4103 ± 0.2026
-0.0620
0.4255 ± 0.2086
Random 10%
0.4511 ± 0.1175
-0.0212
0.4760 ± 0.1127
0.5133 ± 0.0431
0.0409
0.5318 ± 0.0455
Reward 5%
0.4516 ± 0.2104
-0.0208
0.4729 ± 0.2181
0.4443 ± 0.1691
-0.0281
0.4628 ± 0.1733
Table 1: GRPO results on GSM8K under clean and mismatch settings. Δ Acc. is measured relative to the shared pre-trained model. Results are reported as mean ± std. over five seeds, with the best result shown in bold. Given the consistent margins, we omit additional pairwise t -tests.
Figure 3: Empirical analysis of the self- and cross-term effect on GSM8K under Mismatch-20% . We define conflicted units as training units for which DTV and DTV-Loo make different keep/drop decisions; R denotes the fraction of units retained for visualization after trimming extreme score ranges, and IQR denotes the interquartile range (25th–75th percentiles). (a) Decomposition of the DTV score into the self- and cross-terms. (b) Filtering regions of DTV and DTV-Loo. (c) Self- and cross-term scores of conflicted units. (d) Cross-term contribution ∣cj∣/(sj+∣cj∣) for all and conflicted units. (e) Cross-term magnitude of conflicted units. (f) Drop ratios of DTV and the additional filtering induced by DTV-Loo.
Pass@1
Pass@16
Maj@16
Extractable ↑
Acc. given extractable ↑
Method
Score ↑
Δ↑
Score ↑
Δ↑
Score ↑
Δ↑
Pre-trained
0.1063
0.0000
0.3333
0.0000
0.2333
0.0000
0.7292
0.1457
Vanilla GRPO
0.1063
0.0000
0.4000
0.0667
0.2333
0.0000
0.7375
0.1441
DTV (Ours)
0.0917
-0.0146
0.3333
0.0000
0.2667
0.0334
0.7188
0.1275
DTV-Loo (Ours)
0.1333
0.0270
0.4333
0.1000
0.3667
0.1334
0.7521
0.1773
Table 2: GRPO results on AIME 2024 with a 32K maximum generation length. Δ denotes the change relative to the shared pre-trained model. Results are from a single training run for each method, with the best result shown in bold.
Clean
Mismatch-20%
Mismatch-40%
Method
AUC ↑
Acc. ↑
AUC ↑
Acc. ↑
AUC ↑
Acc. ↑
Vanilla DPO
0.5611 ± 0.0035
0.5388 ± 0.0073
0.5377 ± 0.0029
0.5014 ± 0.0044
0.4932 ± 0.0105
0.4301 ± 0.0160
Random 5%
0.5585 ± 0.0022
0.5449 ± 0.0023
0.5371 ± 0.0032
0.5013 ± 0.0152
0.4907 ± 0.0105
0.4234 ± 0.0143
Random 10%
0.5578 ± 0.0017
0.5421 ± 0.0063
0.5363 ± 0.0040
0.5015 ± 0.0100
0.4900 ± 0.0113
0.4332 ± 0.0122
Reward 5%
0.5620 ± 0.0016
0.5559 ± 0.0062
0.5465 ± 0.0028
0.5350 ± 0.0092
0.5083 ± 0.0089
0.4683 ± 0.0191
Reward 10%
0.5652 ± 0.0013
0.5647 ± 0.0063
0.5490 ± 0.0006
0.5443 ± 0.0048
0.5132 ± 0.0090
0.4857 ± 0.0147
Table 3: DPO preference-accuracy results on UltraFeedback under clean and mismatch settings. Results are reported as mean ± std. over five seeds, with the best result in each column shown in bold. † The final row reports two-sided paired t -test p -values comparing DTV-Loo with Reward 10% across matched seeds; significant differences ( p<0.05 ) are shown in bold.
Figure 4: DPO validation dynamics and training efficiency on UltraFeedback . The top row shows the change in validation accuracy relative to Vanilla DPO throughout training under Clean , Mismatch-20% , and Mismatch-40% conditions, averaged over five seeds; shaded regions indicate variability across seeds, and the horizontal dashed line denotes Vanilla DPO. The bottom row reports, for seed 0, the wall-clock time to reach the T95 target defined from the five-seed best validation performance; × NR denotes that the target is not reached within the training budget. DTV-Loo exhibits increasingly large and more persistent improvements as the mismatch rate increases, while also reaching the convergence target substantially faster under the mismatch settings.
Figure 5: Empirical analysis of the self-protection effect on UltraFeedback under Mismatch-20% . (a) Decomposition of the DTV Score into the self- and cross-terms. (b) Cross-term contribution ∣cj∣/(sj+∣cj∣) for all and conflicted units. (c) Visualization of self- and cross-term scores across samples in a representative batch (ranked by self-term). A single dominant sample with an extreme self-term score overwhelms the cross-term interactions across the entire batch, illustrating the self-protection mechanism in standard DTV.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Auxiliary signal
Fixed dataset
Validation set
Selection budget
Extra rollouts
Online valuation
Applied scenario
IIF ( 2025 )
Yes
No
No
Yes
No
Yes
PPO
LearnAlign ( 2026 )
Yes
Yes
No
Yes
Yes
No
GRPO / RLVR
GradAlign ( 2026 )
Yes
No
Yes
Yes
Yes
Yes
GRPO
DTV/DTV-Loo (Ours)
No
No
No
No
No
Yes
General
Appendix
Table 4: Comparison of additional requirements and applicability of gradient-based data valuation methods in reinforcement learning. The binary columns indicate whether each method uses an auxiliary valuation signal, assumes a fixed dataset, requires a validation set, uses an explicit selection budget, requires extra rollouts for valuation, or performs valuation online. The final column summarizes each method’s generalizability across reinforcement learning settings.
Hyperparameter
Empty-8x8
DoorKey-8x8
Parallel environments
16
16
Environment-step budget
160,000
1,000,000
Rollout steps per environment
128
640
Transitions per rollout round
2,048
10,240
Optimizer mini-batch size
64
256
PPO epochs per round
10
4
Appendix
Table 5: PPO training and evaluation configuration. The environment-step budget is the configured stopping target; collection ends after completing the final rollout round.
Hyperparameter
GSM8K
AIME
Prompts per update
4
128
Completions per prompt
4
8
Train micro-batch size
–
2
Maximum prompt / response length
256 / 768
1,024 / 8,192
Optimizer / peak learning rate
AdamW / 1×10−6
AdamW / 1×10−6
Learning-rate schedule
69-step warmup, cosine decay
cosine decay, no warmup
Appendix
Table 6: GRPO training and evaluation configurations on GSM8K and AIME 2024
Hyperparameter
SFT
DPO
Fine-tuning mode
Full parameter
LoRA ( r=64 , scaling =64 )
Train / evaluation batch size
2 / 2
8 / 8
Gradient accumulation steps
4
32
Maximum length
Target: 768
Prompt / response: 512 / 512
Optimizer / peak learning rate
AdamW / 1×10−5
AdamW / 1×10−6
Warmup / decay steps
100 / 1,500
10 / 115
Appendix
Table 7: SFT and DPO training configurations
Condition
Method
LiveBench-IF ↑
RB2 Precise IF ↑
IFBench P-Strict ↑
Clean
Vanilla DPO
0.2418 ± 0.0035
0.4938 ± 0.0285
0.1280 ± 0.0045
Random 5%
0.2413 ± 0.0046
0.4763 ± 0.0294
0.1227 ± 0.0090
Random 10%
0.2436 ± 0.0049
0.4813 ± 0.0442
0.1307 ± 0.0068
Reward 5%
0.2388 ± 0.0064
0.4800 ± 0.0203
0.1253 ± 0.0016
Reward 10%
0.2370 ± 0.0055
0.4613 ± 0.0191
0.1247 ± 0.0050
DTV (Ours)
0.2430 ± 0.0034
0.4938 ± 0.0198
0.1247 ± 0.0045
Appendix
Table 8: Clean downstream instruction-following performance of DPO methods. Results are mean ± sample standard deviation over five training seeds. These metrics are not used for training, filtering, or checkpoint selection; the best result in each condition and column is shown in bold.
Figure 6: Absolute preference-accuracy trajectories on UltraFeedback under Clean , Mismatch-20% , and Mismatch-40% training conditions. Curves show the five-seed mean and shaded regions show ±1 sample standard deviation.
Figure 7: Fixed-filtering-rate ablation under Mismatch-40% . DTV and DTV-Loo are compared with Random and Reward-based Filtering at 5%, 10%, 20%, and 40%. Curves show the five-seed mean, with shading denoting ±1 sample standard deviation.