Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts and uses available metadata to construct group-informed priors for new users. Starting from these fixed priors, the model personalizes independently as feedback from each user becomes available. Session slates combine Thompson sampling with diversity and inventory-depletion controls. We evaluate CohortMix-TS through simulation, semi-synthetic experiments, and a 25-day randomized in-the-wild deployment with 713 registered participants in a Campus Games quiz application. Our evaluations show that cross-cohort transfer improves early recommendation quality and user-level regret, while inventory-aware slate construction helps prevent premature exhaustion of preferred items. In the field deployment, treatment users also showed a larger early-to-late change in correctness than users receiving random recommendations. Together, these results show how warm-start transfer and inventory-aware recommendations can support personalization for short-lived, repeatedly cold-starting cohorts.
Figures & tables
Figure 1. One diagnostic for each generated environment. (a) Transfer expected reward at alignment χ=0.75 for CohortMix-TS , Cold-start TS, Metadata LinUCB, and Metadata LinTS; higher χ means closer agreement between historical and target preferences. (b) Preferred-arm inventory use when that arm contains 20% of the TK item opportunities for No controls, Diversity only, Depletion only, and Full selector. (c) Real-catalog scarcity stress in the 2024–2025 year-split environment. Each line joins five consecutive five-session blocks. Both coordinates average over users whose generated optimal arm is the actual 98-item arm: expected 10-item slate reward (horizontal) and the percentage with an unseen item remaining in that arm before the session slate is selected (vertical). Open markers denote sessions 1–5 and filled terminal markers denote sessions 21–25. Shaded bands in (a)–(b) are percentile bootstrap 95% intervals across 160 cohorts; (c) reports cohort-average trajectories across 96 cohorts. Three panels. The first plots expected reward over 25 sessions for full CohortMix-TS, Cold-start TS, Metadata LinUCB, and Metadata LinTS. The full CohortMix-TS line is highest by the end. The second plots preferred-arm inventory consumed over 25 sessions for four slate-control variants; no controls and diversity only exhaust the arm early, while depletion-aware variants use it gradually. The third plots five connected positions for full CohortMix-TS, Cold-start TS, No depletion, and Random in a real-catalog scarcity stress. Its horizontal axis is expected reward of a ten-item slate for users whose generated optimal arm is a 98-item arm. Its vertical axis is the percentage of those users with an unseen item remaining in that arm before their session slate is selected. An open marker is the first block and a larger filled terminal marker is the fifth.
Policy
Early ↑
Campaign ↑
Minority ↑
P90 regret ↓
Full CohortMix-TS
0.622
0.669
0.653
19.18
Warm TS (fixed mixture)
0.579
0.642
0.628
24.39
Hard-cluster TS
0.592
0.609
0.506
62.77
Cold-start TS
0.576
0.639
0.638
24.73
Static source mixture
0.585
0.585
0.400
87.47
Metadata LinUCB
0.541
0.596
0.505
65.23
Table 1. Parametric transfer results at alignment χ=0.75 , averaged over 160 generated cohorts. Rewards are expected per displayed item; Minority is campaign reward for the lower-probability hidden subtype. P90 pseudo-regret is the 90th percentile of per-user pseudo-regret at the final session. Bold marks the best result and underline the second-best result in each column.
Policy
Early ↑
Campaign ↑
Late ↑
Pseudo-regret ↓
CohortMix-TS
0.597
0.627
0.632
37.71
Cold-start TS
0.584
0.618
0.628
39.94
Metadata LinUCB
0.580
0.590
0.601
46.84
Metadata LinTS
0.580
0.594
0.598
45.99
Hard membership
0.583
0.615
0.625
40.59
Global prior
0.584
0.615
0.624
40.67
Table 2. Year-split data-calibrated semi-synthetic results over 96 generated cohorts. Source-only 2024 structure initializes each policy; 2025 records calibrate hidden target preferences. Reward entries are expected per displayed item; pseudo-regret is per user over 250 displayed items. The main benchmark turns slate penalties off to isolate transfer; Figure 1 (c) reports the separate real-catalog scarcity stress. Bold marks the best result and underline the second-best result in each column.
Figure 2. In-the-wild deployment outcomes. (a) Participant-level early–late daily-correctness change Δu in the restricted, complete-window analysis ( n=52 ; treatment =23 , control =29 ). Points are users; diamonds and bars are group means and bootstrap 95% intervals. The treatment–control difference is 0.062 (95% interval [0.005,0.119] ; p=0.043 ). (b) Daily mean correct/answered rate, averaged within each responding user-day; (c) daily active users; and (d) active-day survival. Panels (b)–(d) are descriptive all-registrant traces ( n=713 ) after removing technical-failure dates and have changing daily denominators. Blue denotes random control and orange denotes CohortMix-TS . Four panels. The first uses violin plots, points, and mean confidence intervals to show early–late daily correctness change for 29 random-control and 23 CohortMix-TS users. The treatment distribution and mean are above the control distribution and mean. The second panel shows descriptive participant-day correct-answer rates, the third the number of daily active users, and the fourth the share of all registrants with at least a given number of active quiz days. Blue represents control and orange represents CohortMix-TS throughout.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Policy
Early ↑
Campaign ↑
Late ↑
Early exhaust ↓
No controls
.4386
.4333
.4320
1.00
Diversity only
.5244
.5002
.4506
1.00
Depletion only
.4756
.4790
.4834
.00
No adaptive controller
.5177
.5081
.4983
.00
Full selector
.5168
.5083
.4996
.00
Appendix
Table 3. Ablation and sensitivity details. (Left) finite-inventory component ablation at m/(TK)=.20 ; rewards are expected per displayed item and Early exhaust is the fraction of users whose preferred arm is exhausted before the final week. “No adaptive controller” retains γ=.5 and δ=1 but sets ϕ=0 . (Right) the early-reward difference between CohortMix-TS and Cold-start TS in the transfer environment; positive values favor CohortMix-TS . Bold marks the best and underline the second-best result within each left-hand reward column.
Figure 3. Source-only group–arm variation conditional on metadata. Within the three highest-support metadata categories that split into multiple inferred groups, each cell is a group’s leave-group-out correctness residual (percentage points) relative to other historical users with the same metadata category and item arm. A–C denote anonymized metadata categories; rows are ordered by source answer support. The gray cell has fewer than 30 observations. A five-column heatmap for historical item arms one through five. Nine rows are grouped into three anonymized metadata categories A, B, and C. Cells show positive or negative leave-group-out correctness residuals in percentage points relative to users with the same metadata category; colors range from blue negative values through white near zero to red positive values. One low-support cell is gray.
Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation. Despite their ability to adapt from online feedback, these methods often suffer from cold-start limitations when newly introduced arms have little or no interaction history. In our setting, the candidate arms are user-generated textual comments, whose semantic content can reveal a title's appeal before sufficient interaction feedback is available. We therefore use large language models (LLMs) to extract semantic signals from comment text and convert them into informative Bayesian priors that warm-start Thompson sampling under sparse early-stage feedback. To account for aggregate segment-level differences in response patterns, we maintain and update posteriors separately for each gender-age segment. In a real-world online A/B/C test, we compare a uniform prior with two LLM-based designs: a Gender Prior for demographic-affinity cues and a Content Prior for title-specific identity cues. The results show that LLM-based priors are most beneficial in sparse-feedback regimes -- with the largest gains emerging once a small amount of interaction evidence has accumulated -- and that prior design leads to distinct funnel-level effects. We further analyze prior-reward alignment and demographic heterogeneity, finding that click-oriented alignment is strongest for the Gender Prior and that treatment effects vary substantially across demographic segments. These findings suggest that LLM-derived priors can serve as a practical warm-start mechanism for text-rich bandit recommendation, while also revealing deployment trade-offs.
Implicit feedback is widely used in recommender systems due to its accessibility and generality, yet it usually presents noisy samples (e.g., clickbait, position bias). Meanwhile, recommenders inevitably face the item cold-start problem due to the continuous influx of new items. We identify that cold items are more prone to noisy samples due to the aforementioned factors, and researchers often overlook the significance of denoising implicit feedback for cold items. Previous denoising studies usually identify noisy samples based on heuristic patterns, such as higher loss values, and mitigate noise through sample selection or re-weighting. However, these methods have limited adaptability and are ineffective in cold-start scenarios. To achieve denoising implicit feedback for cold-start recommendation, we propose a model-agnostic denoising method called DIF. First, user preferences for content remain stable, which allows us to infer pseudo-labels indicating whether a user is interested in a cold item through content-similar warm items. Furthermore, to improve pseudo-label accuracy, we model the confidence of pseudo-labels based on the content similarity between the cold item and warm items, and then aggregate multiple pseudo-labels for each sample. Finally, we explicitly estimate the uncertainty of the noisy sample label by considering its relative entropy and the cold-start status of the item, which adaptively guides the role of pseudo-labels to correct the noisy labels at the sample level. DIF's superiority is supported by both theoretical justification and extensive experiments on real-world datasets. The method has been deployed on a billion-user scale short video application Kuaishou and has significantly improved various commercial metrics within cold-start scenarios.
Gaode Chen, Shicheng Wang, Shikun Li +8
Kuaishou Technology Beijing, China · Hong Kong Baptist University Hong Kong, China · Institute of Information Engineering, Chinese Academy of Sciences Beijing, China
Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration task, which we formalize as a bandit problem with delayed rewards. There is an apparent trade-off in choosing the learning signal: waiting for the full reward to become available might take several weeks, slowing the rate of learning, whereas using short-term proxy rewards reflects the actual long-term goal only imperfectly. First, we develop a predictive model of delayed rewards that incorporates all information obtained to date. Rewards as well as shorter-term surrogate outcomes are combined through a Bayesian filter to obtain a probabilistic belief. Second, we devise a bandit algorithm that quickly learns to identify content aligned with long-term success using this new predictive model. We prove a regret bound for our algorithm that depends on the Value of Progressive Feedback, an information-theoretic metric that captures the quality of short-term leading indicators that are observed prior to the long-term reward. We apply our approach to a podcast recommendation problem, where we seek to recommend shows that users engage with repeatedly over two months. We empirically validate that our approach significantly outperforms methods that optimize for short-term proxies or rely solely on delayed rewards, as demonstrated by an A/B test in a recommendation system that serves hundreds of millions of users.
Kelly W. Zhang, Thomas Baldwin-McDonald, Kamil Ciosek +2
Imperial College London · University of Manchester · Spotify +2