Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching 1.11× to 5.21× that of the strongest baseline, with competitive session depth and no larger retained sets.
Figures & tables
Figure 1: Performance across the 19 experimental configurations. The axes are normalized, and set size is inverted so that larger values are preferable on every axis. RLCP improves catalog diversity while maintaining competitive session depth under the same set size cap as the baselines. The full detailed results and configurations are in Appendix F .
Figure 2: RLCP framework. At state sk , critic scores gθk(sk,a) and threshold τk define the raw set Cθk,τk(sk) . A minimum score fallback ensures nonempty execution. The resulting set constrains the downstream actor in RLCP or supplies the simulator in RLCP Single. Proxy miss feedback updates the threshold more frequently than the critic is updated; this schedule is an implementation choice.
Input: target miss rate α ; interaction budget T ; threshold τ1 ; nonincreasing ηt>0 ; cap K≥1 ; critic Qθ1 ; selector πϕ1 ; proxy-target verifier.
1
for t=1,…,T do
2
Observe st and compute gθt(st,⋅) using ( 20 ).
3
Form Ct={a:gθt(st,a)≤τt} .
4
Keep the min{K,∣Ct∣} lowest-score actions to obtain CtK .
5
Set Dt=CtK if nonempty; otherwise set Dt={aθtmin(st)} .
6
Define the proxy target for the current state and obtain the raw miss et=1{Ct∩Aε,t(st)=∅} .
Algorithm 1 RLCP with Online Action-Set Calibration
Dataset
Users
Items
Interactions
Sessions
Density
KuaiRand -Pure
27,077
7,551
1,436,609
246,738
0.70%
ML-1M
6,400
3,706
1,000,208
16,629
4.22%
Table 1: Statistics of the datasets used.
Figure 3: KuaiRand-Pure radar summaries for the four reward and patience configurations reported in Tables 2 – 5 . Each panel averages over the baseline slate sizes in the corresponding table and shows depth, catalog diversity, intra-list diversity, and inverted set size, with larger radius indicating better performance. RLCP-Single gives the strongest catalog-diversity gains on KuaiRand-Pure while using equal or smaller admissible sets, whereas the full RLCP consistently attains the highest intra-list diversity.
Figure 4: ML-1M radar summaries for the four reward and patience configurations reported in Tables 6 – 9 . Each panel averages over the baseline slate sizes in the corresponding table and uses the same normalized axes as Figure 3 . On the denser ML-1M data, the full RLCP is more competitive in terms of depth and catalog diversity while maintaining maximal intra-list diversity, and the RLCP variants retain their overall diversity advantage over the fixed-slate baselines.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=3
DDPG
18.80
0.62
13.46
25.00
0.99
3.00
0.62
0.61
0.54
0.54
0.55
0.58
0.64
TD3
19.20
0.65
13.24
54.50
0.99
3.00
0.65
0.62
0.52
0.54
0.52
0.58
0.63
A2C
19.60
0.66
13.32
3.00
0.99
3.00
0.66
0.61
0.58
0.56
0.54
0.60
0.62
HAC
19.10
0.67
13.47
4.70
0.98
3.00
0.67
0.62
0.57
0.55
0.54
0.56
0.68
RLCP-Single
19.30
0.64
13.20
92.40
0.99
2.91
0.64
0.57
0.55
0.55
0.52
0.59
0.65
Appendix
Table 2: KuaiRand-Pure results under patience setting (p,q)=(1,2) with is_click as the reward signal for displayed slate sizes M∈{3,4,5} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=4
DDPG
46.80
0.70
32.83
14.40
0.99
4.00
0.70
0.62
0.59
0.57
0.52
0.60
0.68
TD3
48.00
0.69
33.46
25.90
0.99
4.00
0.69
0.64
0.59
0.60
0.53
0.61
0.70
A2C
46.00
0.64
31.35
4.10
0.98
4.00
0.64
0.60
0.57
0.54
0.55
0.57
0.63
HAC
48.00
0.68
32.95
7.10
0.99
4.00
0.68
0.64
0.58
0.57
0.53
0.61
0.69
RLCP-Single
47.60
0.68
32.92
118.10
0.99
3.84
0.68
0.62
0.57
0.54
0.55
0.60
0.65
Appendix
Table 3: KuaiRand-Pure results under patience setting (p,q)=(0.4,2) with is_click as the reward signal for displayed slate sizes M∈{4,5} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=4
DDPG
19.80
0.61
12.79
24.00
0.99
4.00
0.66
0.60
0.58
0.55
0.54
0.61
0.65
TD3
19.80
0.60
12.69
51.80
0.99
4.00
0.62
0.60
0.59
0.57
0.56
0.60
0.62
A2C
19.70
0.60
12.83
4.00
0.98
4.00
0.62
0.60
0.59
0.55
0.53
0.60
0.59
HAC
20.00
0.58
12.84
8.60
0.99
4.00
0.64
0.59
0.56
0.57
0.53
0.58
0.66
RLCP-Single
19.80
0.61
12.85
119.70
0.99
3.91
0.63
0.59
0.57
0.57
0.54
0.61
0.63
Appendix
Table 4: KuaiRand-Pure results under patience setting (p,q)=(1,2) with is_like as the reward signal for displayed slate sizes M∈{4,5,6} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Comment
Follow
Forward
Hate
Like
Long view
Slate size M=5
DDPG
47.60
0.57
31.62
35.20
0.99
5.00
0.65
0.60
0.57
0.55
0.57
0.57
0.67
TD3
48.00
0.60
31.93
53.90
0.99
5.00
0.69
0.64
0.57
0.58
0.56
0.60
0.71
A2C
45.60
0.60
30.79
5.00
0.98
5.00
0.64
0.62
0.58
0.56
0.55
0.60
0.66
HAC
47.20
0.61
29.78
5.10
0.98
5.00
0.63
0.59
0.59
0.52
0.54
0.61
0.63
RLCP-Single
48.00
0.62
31.09
151.70
0.99
4.84
0.70
0.63
0.61
0.57
0.55
0.62
0.70
Appendix
Table 5: KuaiRand-Pure results under patience setting (p,q)=(0.4,2) with is_like as the reward signal for displayed slate sizes M∈{5,6} . The final seven columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
19.90
0.69
13.95
23.90
0.99
4.00
0.69
0.69
0.70
TD3
19.90
0.73
14.90
15.90
0.99
4.00
0.73
0.72
0.70
A2C
19.20
0.58
11.97
46.30
0.99
4.00
0.58
0.61
0.62
HAC
20.00
0.68
14.30
4.80
0.99
4.00
0.68
0.66
0.66
RLCP-Single
19.80
0.58
11.85
119.40
0.99
4.00
0.58
0.58
0.58
Appendix
Table 6: ML-1M results under patience setting (p,q)=(1,2) with is_click as the reward signal for displayed slate sizes M∈{4,5} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
46.80
0.69
33.76
12.90
0.99
4.00
0.69
0.71
0.73
TD3
46.80
0.71
35.13
14.30
0.99
4.00
0.71
0.71
0.73
A2C
44.00
0.59
26.76
31.20
0.99
4.00
0.59
0.64
0.64
HAC
45.20
0.61
29.02
9.70
0.99
4.00
0.61
0.64
0.65
RLCP-Single
45.60
0.57
26.56
119.30
0.99
3.97
0.57
0.59
0.63
Appendix
Table 7: ML-1M results under patience setting (p,q)=(0.4,2) with is_click as the reward signal for displayed slate sizes M∈{4,5} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
19.80
0.65
13.21
32.80
0.99
4.00
0.63
0.65
0.67
TD3
19.80
0.66
13.71
20.80
0.99
4.00
0.65
0.66
0.66
A2C
19.60
0.60
12.83
4.10
0.99
4.00
0.57
0.60
0.60
HAC
19.90
0.64
13.31
13.90
0.99
4.00
0.67
0.64
0.63
RLCP-Single
19.80
0.63
12.75
7.70
0.99
4.00
0.58
0.63
0.62
Appendix
Table 8: ML-1M results under patience setting (p,q)=(1,2) with is_like as the reward signal for displayed slate sizes M∈{4,5,6} . The final three columns report behavior rates; set size is reported for RLCP variants when available.
Method
Depth
Avg. reward
Total reward
Diversity
ILD
Set size
Click
Like
Star
Slate size M=4
DDPG
47.20
0.73
33.86
20.10
0.99
4.00
0.68
0.73
0.70
TD3
47.20
0.70
34.57
20.60
0.99
4.00
0.69
0.70
0.70
A2C
41.20
0.64
29.45
22.30
0.99
4.00
0.60
0.64
0.63
HAC
44.20
0.71
31.59
6.80
0.99
4.00
0.68
0.71
0.69
RLCP-Single
46.80
0.63
29.46
12.40
0.99
4.00
0.61
0.63
0.63
Appendix
Table 9: ML-1M results under patience setting (p,q)=(0.4,2) with is_like as the reward signal for displayed slate sizes M∈{4,5} . The final three columns report behavior rates; set size is reported for RLCP variants when available.