Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching 1.11× to 5.21× that of the strongest baseline, with competitive session depth and no larger retained sets.
Figures & tables
Figure 1: Performance across the 19 experimental configurations. The axes are normalized, and set size is inverted so that larger values are preferable on every axis. RLCP improves catalog diversity while maintaining competitive session depth under the same set size cap as the baselines. The full detailed results and configurations are in Appendix F .
Figure 2: RLCP framework. At state sk , critic scores gθk(sk,a) and threshold τk define the raw set Cθk,τk(sk) . A minimum score fallback ensures nonempty execution. The resulting set constrains the downstream actor in RLCP or supplies the simulator in RLCP Single. Proxy miss feedback updates the threshold more frequently than the critic is updated; this schedule is an implementation choice.
Input: target miss rate α ; interaction budget T ; threshold τ1 ; nonincreasing ηt>0 ; cap K≥1 ; critic Qθ1 ; selector πϕ1 ; proxy-target verifier.
1
for t=1,…,T do
2
Observe st and compute gθt(st,⋅) using ( 20 ).
3
Form Ct={a:gθt(st,a)≤τt} .
4
Keep the min{K,∣Ct∣} lowest-score actions to obtain CtK .
5
Set Dt=CtK if nonempty; otherwise set Dt={aθtmin(st)} .
6
Define the proxy target for the current state and obtain the raw miss et=1{Ct∩Aε,t(st)=∅} .
Algorithm 1 RLCP with Online Action-Set Calibration
Dataset
Users
Items
Interactions
Sessions
Density
KuaiRand -Pure
27,077
7,551
1,436,609
246,738
0.70%
ML-1M
6,400
3,706
1,000,208
16,629
4.22%
Table 1: Statistics of the datasets used.
Figure 3: KuaiRand-Pure radar summaries for the four reward and patience configurations reported in Tables 2 – 5 . Each panel averages over the baseline slate sizes in the corresponding table and shows depth, catalog diversity, intra-list diversity, and inverted set size, with larger radius indicating better performance. RLCP-Single gives the strongest catalog-diversity gains on KuaiRand-Pure while using equal or smaller admissible sets, whereas the full RLCP consistently attains the highest intra-list diversity.
Figure 4: ML-1M radar summaries for the four reward and patience configurations reported in Tables 6 – 9 . Each panel averages over the baseline slate sizes in the corresponding table and uses the same normalized axes as Figure 3 . On the denser ML-1M data, the full RLCP is more competitive in terms of depth and catalog diversity while maintaining maximal intra-list diversity, and the RLCP variants retain their overall diversity advantage over the fixed-slate baselines.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=3
DDPG
18.80
0.62
13.46
25.00
0.99
3.00
0.62
0.61
0.54
0.54
0.55
0.58
0.64
TD3
19.20
0.65
13.24
54.50
0.99
3.00
0.65
0.62
0.52
0.54
0.52
0.58
0.63
A2C
19.60
0.66
13.32
3.00
0.99
3.00
0.66
0.61
0.58
0.56
0.54
0.60
0.62
HAC
19.10
0.67
13.47
4.70
0.98
3.00
0.67
0.62
0.57
0.55
0.54
0.56
0.68
RLCP-Single
19.30
0.64
13.20
92.40
0.99
2.91
0.64
0.57
0.55
0.55
0.52
0.59
0.65
Appendix
Table 2: KuaiRand-Pure results under patience setting (p,q)=(1,2) with is_click as the reward signal for displayed slate sizes M∈{3,4,5} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=4
DDPG
46.80
0.70
32.83
14.40
0.99
4.00
0.70
0.62
0.59
0.57
0.52
0.60
0.68
TD3
48.00
0.69
33.46
25.90
0.99
4.00
0.69
0.64
0.59
0.60
0.53
0.61
0.70
A2C
46.00
0.64
31.35
4.10
0.98
4.00
0.64
0.60
0.57
0.54
0.55
0.57
0.63
HAC
48.00
0.68
32.95
7.10
0.99
4.00
0.68
0.64
0.58
0.57
0.53
0.61
0.69
RLCP-Single
47.60
0.68
32.92
118.10
0.99
3.84
0.68
0.62
0.57
0.54
0.55
0.60
0.65
Appendix
Table 3: KuaiRand-Pure results under patience setting (p,q)=(0.4,2) with is_click as the reward signal for displayed slate sizes M∈{4,5} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=4
DDPG
19.80
0.61
12.79
24.00
0.99
4.00
0.66
0.60
0.58
0.55
0.54
0.61
0.65
TD3
19.80
0.60
12.69
51.80
0.99
4.00
0.62
0.60
0.59
0.57
0.56
0.60
0.62
A2C
19.70
0.60
12.83
4.00
0.98
4.00
0.62
0.60
0.59
0.55
0.53
0.60
0.59
HAC
20.00
0.58
12.84
8.60
0.99
4.00
0.64
0.59
0.56
0.57
0.53
0.58
0.66
RLCP-Single
19.80
0.61
12.85
119.70
0.99
3.91
0.63
0.59
0.57
0.57
0.54
0.61
0.63
Appendix
Table 4: KuaiRand-Pure results under patience setting (p,q)=(1,2) with is_like as the reward signal for displayed slate sizes M∈{4,5,6} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=5
DDPG
47.60
0.57
31.62
35.20
0.99
5.00
0.65
0.60
0.57
0.55
0.57
0.57
0.67
TD3
48.00
0.60
31.93
53.90
0.99
5.00
0.69
0.64
0.57
0.58
0.56
0.60
0.71
A2C
45.60
0.60
30.79
5.00
0.98
5.00
0.64
0.62
0.58
0.56
0.55
0.60
0.66
HAC
47.20
0.61
29.78
5.10
0.98
5.00
0.63
0.59
0.59
0.52
0.54
0.61
0.63
RLCP-Single
48.00
0.62
31.09
151.70
0.99
4.84
0.70
0.63
0.61
0.57
0.55
0.62
0.70
Appendix
Table 5: KuaiRand-Pure results under patience setting (p,q)=(0.4,2) with is_like as the reward signal for displayed slate sizes M∈{5,6} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
19.90
0.69
13.95
23.90
0.99
4.00
0.69
0.69
0.70
TD3
19.90
0.73
14.90
15.90
0.99
4.00
0.73
0.72
0.70
A2C
19.20
0.58
11.97
46.30
0.99
4.00
0.58
0.61
0.62
HAC
20.00
0.68
14.30
4.80
0.99
4.00
0.68
0.66
0.66
RLCP-Single
19.80
0.58
11.85
119.40
0.99
4.00
0.58
0.58
0.58
Appendix
Table 6: ML-1M results under patience setting (p,q)=(1,2) with is_click as the reward signal for displayed slate sizes M∈{4,5} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
46.80
0.69
33.76
12.90
0.99
4.00
0.69
0.71
0.73
TD3
46.80
0.71
35.13
14.30
0.99
4.00
0.71
0.71
0.73
A2C
44.00
0.59
26.76
31.20
0.99
4.00
0.59
0.64
0.64
HAC
45.20
0.61
29.02
9.70
0.99
4.00
0.61
0.64
0.65
RLCP-Single
45.60
0.57
26.56
119.30
0.99
3.97
0.57
0.59
0.63
Appendix
Table 7: ML-1M results under patience setting (p,q)=(0.4,2) with is_click as the reward signal for displayed slate sizes M∈{4,5} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
19.80
0.65
13.21
32.80
0.99
4.00
0.63
0.65
0.67
TD3
19.80
0.66
13.71
20.80
0.99
4.00
0.65
0.66
0.66
A2C
19.60
0.60
12.83
4.10
0.99
4.00
0.57
0.60
0.60
HAC
19.90
0.64
13.31
13.90
0.99
4.00
0.67
0.64
0.63
RLCP-Single
19.80
0.63
12.75
7.70
0.99
4.00
0.58
0.63
0.62
Appendix
Table 8: ML-1M results under patience setting (p,q)=(1,2) with is_like as the reward signal for displayed slate sizes M∈{4,5,6} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
47.20
0.73
33.86
20.10
0.99
4.00
0.68
0.73
0.70
TD3
47.20
0.70
34.57
20.60
0.99
4.00
0.69
0.70
0.70
A2C
41.20
0.64
29.45
22.30
0.99
4.00
0.60
0.64
0.63
HAC
44.20
0.71
31.59
6.80
0.99
4.00
0.68
0.71
0.69
RLCP-Single
46.80
0.63
29.46
12.40
0.99
4.00
0.61
0.63
0.63
Appendix
Table 9: ML-1M results under patience setting (p,q)=(0.4,2) with is_like as the reward signal for displayed slate sizes M∈{4,5} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinforcement learning (RL) provides a principled framework for optimizing such sequential decision tasks, as path rewards can naturally capture both short-term acceptance and long-term guidance effectiveness. However, naively applying policy gradients to PRS results in deficient gradient estimation. We identify two deficiencies: (1) path-level rewards decompose into step-level rewards with positive mean, creating a length-dependent bias that causes gradients to favor path extension over meaningful exploration; (2) weighting each step by the entire path-level reward ignores the decomposition structure, leading to high gradient variance. To rectify these two deficiencies, we propose an effective RL framework ProRL with two novel mechanisms for proactive recommendation. First, Stepwise Reward Centering subtracts expected rewards to neutralize length-dependent bias, ensuring that path extension yields zero expected gradient signal. Second, Position-Specific Advantage Estimation leverages the reward decomposition structure to compute step-dependent baselines, reducing gradient variance. Together, these mechanisms yield policy gradients that precisely target path quality. Our experiments on three real-world datasets demonstrate that ProRL significantly outperforms state-of-the-art PRSs. Our code is available at https://github.com/hongruhou89/ProRL.
Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refinement, relying on increasingly deep architectures to capture complex patterns. In this work, we revisit sequential recommendation from a recursive inference perspective: can user preferences be modeled as a persistent latent state that is recursively refined? We propose RecRec (Recursive Recommendation), a lightweight model that maintains a compact latent state and updates it through a shared recursive module conditioned on interaction evidence. Unlike prior recursive models, RecRec introduces an evidence-anchored correction mechanism that stabilizes refinement by grounding each update in the original interaction context, preventing semantic drift during deep recursive reasoning. Experiments on three benchmark datasets under standard evaluation protocols show that RecRec matches or outperforms state-of-the-art sequential, graph-based, and reasoning-enhanced recommenders while using only 3.9M to 14M parameters. Ablation studies demonstrate that both recursive refinement and the evidence-anchored correction gate contribute significantly to performance, highlighting the effectiveness of recursive latent inference as a scalable alternative to deeper or language-based architectures. Code is available at https://anonymous.4open.science/r/RecRec-6B67/README.md.
Generic group-based RL assumes that sampled rollout groups are already usable learning signals. We show that this assumption breaks down in sparse-hit generative recommendation, where many sampled groups never become learnable at all. We propose ReCast, a repair-then-contrast learning-signal framework that first restores minimal learnability for all-zero groups and then replaces full-group reward normalization with a boundary-focused contrastive update on the strongest positive and the hardest negative. ReCast leaves the outer RL framework unchanged, modifies only within-group signal construction, and partially decouples rollout search width from actor-side update width. Across multiple generative recommendation tasks, ReCast consistently outperforms OpenOneRec-RL, achieving up to 36.6% relative improvement in Pass@1. Its matched-budget advantage is substantially larger: ReCast reaches the baseline's target performance with only 4.1% of the rollout budget, and this advantage widens with model scale. The same design also yields direct system-level gains, reducing actor-side update time by 16.60x, lowering peak allocated memory by 16.5%, and improving actor MFU by 14.2%. Mechanism analysis shows that ReCast mitigates the persistent all-zero / single-hit regime, restores learnability when natural positives are scarce, and converts otherwise wasted rollout budget into more stable policy updates. These results suggest that, for generative recommendation, the decisive RL problem is not only how to assign rewards, but how to construct learnable optimization events from sparse, structured supervision.