Decoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores
Organizations: University of Bergen Centre for the Science of Learning & Technology (SLATE) · City University of Macau School of Education
Abstract
We formalize slate recommendation as a randomized score learner followed by deterministic selection. First, an appropriately scoped differential-privacy guarantee passes through selection and its audit trace by post-processing. End-to-end privacy holds only when selector inputs are public or independent, previous private outputs, or separately privacy-accounted; fixing raw state or candidate information instead yields only a conditional guarantee. Second, we derive a logged margin certificate: bounded score-induced objective movement below half the smallest greedy decision margin guarantees that the ordered slate is unchanged. Controlled fixed-margin tests show near-linear exponent scaling, with an empirical slope of (95% CI ) against the independent-noise reference . Real-anchor experiments on OULAD, MovieLens-25M, and Amazon Musical Instruments show that greater anchor weight reduces score-noise-induced ranking churn. OULAD and EdNet certificate checks validate the implementation of the logged inequality, while closed-loop simulations show bounded target drift and setting-dependent downstream utility. The contribution is therefore a privacy-scope contract and a certifiable score-to-slate stability mechanism, not a universal utility claim.
Figures & tables
| Quantity | Plain meaning |
|---|---|
| or | The smallest logged decision margin: how clearly the chosen item beat the closest runner-up. |
| or | The largest score-induced movement in the selector objective. |
| A certificate that the ordered slate is unchanged. | |
| The gap between anchor top- items and the remaining candidates. | |
| Anchor weight; larger gives the noisy adaptive score less influence. |
| Claim being tested | Diagnostic | Strongest result |
|---|---|---|
| Top- exponent | Fixed-margin calibration under independent and low-rank correlated score noise | Independent noise: slope (95% CI ) vs. reference ( ). Correlated noise: slope vs. reference ( ). |
| Anchoring reduces raw ranking churn | Real-anchor sweeps on OULAD, MovieLens-25M, and Amazon Musical Instruments | Anchor weight cuts Flip@ by 53–81% across the three corpora and tested noise scales; rating-threshold sensitivity preserves the effect. |
| Pairwise noise condition | Injected-noise tails and descriptive 200-seed learner-update diagnostic | Normalized cross-boundary score differences stay below the sub-Gaussian envelope; shared seeds, pools, and scores make the repeated pair draws descriptive rather than independent observations. |
| Score-to-slate bridge | Ex post certificate-implementation checks on OULAD and EdNet | Anchor gaps and greedy margins are positive in all 640 logged pools. With candidates, state, bonuses, and novelty fixed, the realized check has zero instrumentation violations across 16,000 trials. |
| Target-value drift | Realized-window SCPO check and soft-window event logging | SCPO uses the exact realized-window envelope; for soft-window MMR, the fixed-window event holds on more than 98% of logged round pairs. |
| Downstream stress test | Nominal OULAD update configurations and explicit EdNet score-noise injection | The interpretable noise-injection result gives clean/noisy Kendall- of 0.928–0.993 anchored vs. 0.748–0.904 unanchored; utility changes remain mixed. |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| OULAD | EdNet | |
|---|---|---|
| # Users | 23,351 | 296,701 |
| # Items | 188 | 11,555 |
| # Interactions | 173,739 | 23,384,480 |
| Density (‰) | 39.58 | 6.82 |
| Avg. interactions per user | 7.44 | 78.8 |
| Setting | Standard | Strong | Locked |
|---|---|---|---|
| Approx. ( ) | |||
| Approx. ( ) | |||
| Clipping norm | 1.0 | 1.0 | 1.0 |
| Noise multiplier | 1.2 | 2.0 | 3.0 |
| Sampling rate | 0.020 | 0.015 | 0.010 |
| Family | Primary object usually analyzed | Usually outside the primary analysis | Relation to this paper |
| DP recommendation ( McSherry and Mironov, 2009 ; Abadi et al., 2016 ; Dwork and Roth, 2014 ) | Privacy of the learned model or training algorithm | Deterministic post-layer, audit trace, and score-to-slate stability | Private regime lifts learner DP through a deterministic selector only under admissible/accounted side inputs |
| Re-ranking / slate construction ( Carbonell and Goldstein, 1998 ; Kunaver and Požrl, 2017 ; Steck, 2018 ; Singh and Joachims, 2018 ; Geyik et al., 2019 ) | Deterministic slate objective for diversity, calibration, fairness, or novelty | Learner privacy and explicit scorer-noise model | This paper treats the selector as a bounded surface and studies how it composes with learner privacy and score stability |
| Contextual & private bandits ( Li et al., 2010 ; Chu and Li, 2011 ; Abbasi-Yadkori et al., 2011 ; Shariff and Sheffet, 2018 ; Mishra and Thakurta, 2015 ) | Exploration/regret, sometimes with privacy built into the bandit | Replayable deterministic selection layer and shaped/diverse slate assembly | Proposition 10 recovers LinUCB only in the reduced exploration-only regime |
| Safe / constrained RL for recs ( Ie et al., 2019 ; Achiam et al., 2017 ) | Constraint satisfaction inside a learned policy | Privacy composition and deterministic audit replay | Lemma 4 gives a selector-side drift bound under window enforcement without retraining the learner |
| Off-policy learning / evaluation ( Swaminathan and Joachims, 2015 ) | Estimating or optimizing policy value from logged data | Structural privacy/audit/stability guarantees of the deployed post-layer | Complementary: such estimators could train before the deterministic selector is applied |
| Governance / responsible AI ( Slade and Prinsloo, 2013 ; Drachsler and Greller, 2016 ; Mitchell et al., 2019 ) | Documentation, review, and process artifacts | Formal algorithmic replay and privacy accounting | The audit trace is a formal decision artifact that governance processes could inspect |
| Quantity | Mean | Median | Positive rate | Near-zero rate |
|---|---|---|---|---|
| 0.08462 | 0.05775 | 1.000 | 0.000 | |
| 0.0008477 | 0.0007332 | 1.000 | 0.000 |
| Certified | Same order | Same set | Certified violations | Median | |
|---|---|---|---|---|---|
| 5e-05 | 0.9962 | 1 | 1 | 0 | 0.1343 |
| 0.0001 | 0.9844 | 1 | 1 | 0 | 0.2698 |
| 0.0002 | 0.4306 | 0.9975 | 0.9988 | 0 | 0.537 |
| 0.0005 | 0.01188 | 0.9125 | 0.9975 | 0 | 1.347 |
| 0.001 | 0 | 0.6569 | 0.9875 | 0 | 2.725 |
| Quantity | Mean | Median | Positive rate | Near-zero rate |
|---|---|---|---|---|
| 0.01007 | 0.00706 | 1.000 | 0.000 | |
| 0.0003949 | 0.0002794 | 1.000 | 0.000 |
| Certified | Same order | Same set | Certified violations | Median | |
|---|---|---|---|---|---|
| 5e-05 | 0.5931 | 0.9519 | 0.9775 | 0 | 0.3442 |
| 0.0001 | 0.3944 | 0.8925 | 0.9587 | 0 | 0.682 |
| 0.0002 | 0.1856 | 0.8275 | 0.9375 | 0 | 1.387 |
| 0.0005 | 0.003125 | 0.6669 | 0.8962 | 0 | 3.452 |
| 0.001 | 0 | 0.4275 | 0.8225 | 0 | 7.198 |