Predicting continuous values such as watch-time and gross merchandise value (GMV) is a core problem in industrial recommendation systems. Its inherent difficulty stems from the highly complex and long-tailed distributions of the target signals, which are hard to model accurately. Existing regression methods typically rely on fixed parametric assumptions on the target distribution: overly simple assumptions underfit real-world data, whereas more intricate ones tend to sacrifice scalability and generalization. To address these limitations, we propose a sequence modeling framework based on residual quantization (RQ), in which the target continuous value is decomposed into a sequence of quantization codes that represent progressively finer approximations. The model autoregressively predicts these codes from coarse to fine granularity, with each step refining the residual error left by the previous one. To further improve the quality of the learned representations, we introduce an ordinal-aware representation learning objective that aligns the RQ code embedding space with the ordinal structure of target values, thereby yielding continuous representations of quantization codes and more accurate predictions. We conduct comprehensive experiments on public benchmarks for watch-time and lifetime value (LTV) prediction, together with a large-scale online A/B test for GMV prediction on an industrial short-video recommendation platform. Across all settings, the proposed method shows competitive performance among existing state-of-the-art approaches and generalizes well across diverse continuous value prediction scenarios.
Figures & tables
Figure 1. GMV distribution in a large-scale online short-video commercial scenario. From top to bottom, each subfigure shows the GMV frequency distribution over the value ranges 10–20, 20–30, 30–40, and 40–50.
Figure 2. Architecture of the proposed RQ-Reg method.
Method
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
WLR ( Covington et al., 2016 )
6.047
0.525
0.998
0.672
D2Q ( Zhan et al., 2022 )
5.426
0.565
0.899
0.661
TPM ( Lin et al., 2023 )
4.741
0.599
0.884
0.676
CREAD ( Sun et al., 2024 )
3.215
0.601
0.865
0.678
SWaT ( Yang et al., 2025 )
3.363
0.609
0.847
0.683
Table 1. Performance comparison on watch-time prediction datasets
Dataset
Method
MAE ↓
Gini (+) ↑
SRCC (+) ↑
Gini ↑
SRCC ↑
Criteo-SSC
Two-stage ( Drachen et al., 2018 )
21.719
0.2204
0.2565
0.5278
0.2386
MTL-MSE ( Ma et al., 2018 )
21.190
0.4340
0.3663
0.6330
0.2478
ZILN ( Wang et al., 2019 )
20.880
0.4426
0.3874
0.6338
0.2434
MDME ( Li et al., 2022 )
16.598
0.2297
0.2952
0.4383
0.2269
MDAN ( Liu et al., 2024 )
20.030
0.4128
0.3521
0.6209
0.2470
OptDist ( Weng et al., 2024 )
15.784
0.4428
0.3903
0.6437
0.2505
Table 2. Performance comparison on LTV prediction datasets
Method
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
RQ-Reg
3.189
0.616
0.811
0.695
w/o Lgen
3.199
0.613
0.821
0.691
w/o Lrnc
3.195
0.614
0.819
0.692
w/o ferr
3.191
0.614
0.813
0.694
w/o freg + ferr
3.200
0.613
0.815
0.694
Table 3. Ablation study on watch-time prediction datasets
Method
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
RNN
3.210
0.613
0.815
0.694
LSTM
3.189
0.616
0.811
0.695
Transformer
3.204
0.614
0.829
0.694
Table 4. Comparison of different model backbones on watch-time prediction datasets
Method
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
K-means
3.196
0.611
0.823
0.692
RQ K-means
3.189
0.616
0.811
0.695
Dynamic quantile ( Ma et al., 2024 )
3.242
0.611
0.819
0.693
Table 5. Comparison of different quantization methods on watch-time prediction datasets
Method
Reconst. error
Code length
K-means
7.51×10−2
1
RQ K-means
1.33×10−3
3
Dynamic quantile ( Ma et al., 2024 )
3.28×10−3
3.21 (on avg.)
Table 6. Reconstruction errors (in MAE) of different quantization methods on KuaiRec
Figure 3. Results of RQ K-means quantization error (a) and RQ-Reg prediction (b) with different numbers of RQ levels L on KuaiRec.
Figure 4. Learned representations of target values from different models on KuaiRec. Point color represents the magnitude of the corresponding target value, from low (cool) to high (warm).
Setting
AUC ↑
ADVV ↑
Overall
+0.12%
+4.19%
Long-tail
+0.23%
+4.76%
Table 7. Performance lift of RQ-Reg over the production baseline in the online short-video advertising scenario
Figure 5. Daily ADVV gains during the online evaluation. The A/B test period spans days 1–7. On day 8 the experimental bucket was switched back to the production baseline, and the subsequent A/A test spans days 9–13.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
K=16
3.195
0.614
0.820
0.692
K=32
3.191
0.614
0.818
0.694
K=48
3.189
0.616
0.811
0.695
K=64
3.194
0.614
-
-
Appendix
Table 8. Performance comparison with different settings of hyperparameter K
Configuration
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
α=0.1
3.197
0.614
0.816
0.693
α=0.5
3.189
0.614
0.812
0.695
α=1.0
3.189
0.616
0.811
0.695
α=2.0
3.193
0.614
0.814
0.693
α=5.0
3.192
0.614
0.819
0.692
Appendix
Table 9. Performance comparison with different settings of hyperparameters α and β
Configuration
KuaiRec
CIKM16
MAE ↓
XAUC ↑
MAE ↓
XAUC ↑
k=0.01
3.193
0.614
0.817
0.693
k=0.05
3.191
0.615
0.816
0.694
k=0.1
3.189
0.616
0.811
0.695
k=0.2
3.188
0.615
0.818
0.692
k=0.5
3.193
0.615
0.819
0.692
Appendix
Table 10. Performance comparison with different settings of hyperparameters k and t0
Watch time has emerged as a pivotal metric for optimizing deep user engagement in short-video recommender systems. However, current methods of watch time prediction (WTP) suffer from inherent paradigm-specific limitations. Direct Regression faces mean-collapse due to unimodal Gaussian assumptions, while Ordinal Regression is hampered by quantization errors from rigid discretization. Similarly, Discrete Generative Regression struggles with high inference latency and heuristic vocabulary design. Beyond these specific flaws, a shared deficiency is the inability to capture the intrinsic multimodality and heterogeneity of User-Item Interaction Patterns. To address these challenges, we first revisit the WTP problem from a causal perspective and identify these user-specific patterns as structural confounders that modulate watch time outcomes, where identical interests manifest as distinct watch time outcomes conditioned on diverse user habits. Then, we formally propose a new (or the fourth) paradigm -- Continuous Generative Regression, and introduce FlowTime, a novel method utilizing a One-step Generative Variational Autoencoder. FlowTime effectively circumvents the latency of iterative denoising while maintaining the expressivity of continuous latent spaces. Furthermore, we design a Flow-based Personalized Prior that leverages NFs to warp a standard Gaussian prior into a complex, history-conditioned manifold, thereby enabling the adaptive modeling of multimodal interaction patterns. Finally, we build TimeRec, the first open-source WTP Library, alongside a novel personalization metric to establish a rigorous benchmarking standard. Extensive offline experiments and online A/B tests demonstrate FlowTime's significant superiority over SOTA methods.
Hongxu Ma, Han Zhou, Chenghou Jin +5
Fudan University Shanghai, China · Kuaishou Technology Beijing, China · Shanghai University of Finance and Economics Shanghai, China +1
Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficult to systematically enhance their recommendation capabilities. To address this problem, we propose Probabilistic Residual Learning (PRL), a causal Bayesian recommendation model that models the residual between ground-truth and base predictions, enabling targeted refinement of existing systems. Specifically, PRL (1) probabilistically groups users for localized residual modeling, (2) models domain-level confounders that influence user and item representations, and (3) aggregates cluster-specific residual predictions over the confounders using do-calculus. Experiments demonstrate that our plug-and-play PRL is compatible with various base deep learning recommender systems, improving their performance while automatically discovering meaningful user clusters.
Wenyuan Wang, Yusong Zhao, Zihao Xu +11
Rutgers University · Piscataway, New Jersey, United States · Meta +4
Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refinement, relying on increasingly deep architectures to capture complex patterns. In this work, we revisit sequential recommendation from a recursive inference perspective: can user preferences be modeled as a persistent latent state that is recursively refined? We propose RecRec (Recursive Recommendation), a lightweight model that maintains a compact latent state and updates it through a shared recursive module conditioned on interaction evidence. Unlike prior recursive models, RecRec introduces an evidence-anchored correction mechanism that stabilizes refinement by grounding each update in the original interaction context, preventing semantic drift during deep recursive reasoning. Experiments on three benchmark datasets under standard evaluation protocols show that RecRec matches or outperforms state-of-the-art sequential, graph-based, and reasoning-enhanced recommenders while using only 3.9M to 14M parameters. Ablation studies demonstrate that both recursive refinement and the evidence-anchored correction gate contribute significantly to performance, highlighting the effectiveness of recursive latent inference as a scalable alternative to deeper or language-based architectures. Code is available at https://anonymous.4open.science/r/RecRec-6B67/README.md.