Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.
Figures & tables
Avazu
Criteo
ML-1M
KDD12
KKBox
Models
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
DNN Covington et al. (2016)
0.3721
79.27
0.4380
81.40
0.3100
90.30
0.1502
80.52
0.4811
85.01
PNN Qu et al. (2016)
0.3712
79.44
0.4378
81.42
0.3070
90.42
0.1504
80.47
0.4793
85.15
Wide & Deep Cheng et al. (2016)
0.3720
79.29
0.4376
81.42
0.3056
90.45
0.1504
80.48
0.4852
85.04
DeepFM Guo et al. (2017)
0.3719
79.30
0.4375
81.43
0.3073
90.51
0.1501
80.60
0.4785
85.31
DCNv1 Wang et al. (2017)
0.3719
79.31
0.4376
81.44
0.3156
90.38
0.1501
80.59
0.4766
85.31
Table 1: Performance comparison of different deep CTR models. Typically, CTR researchers consider an improvement of 0.001 (0.1%) in Logloss and AUC to be practically meaningful Zhu et al. (2021) ; Li et al. (2026) . We conduct a two-tailed t-test over five independent runs to assess the statistical significance between our models and the best baseline (*: p<0.01). Best results are highlighted in bold , second-best are underlined .
Model
EMA
KD
Routes
AUC
Recursive backbone
–
–
–
81.59
+ trajectory scaling
✓
–
–
81.65
+ cross-run scaling
–
✓
–
81.68
+ route scaling
–
–
✓
81.61
RECAP
✓
✓
✓
81.76
Table 2: Ablation of estimator-scaling components on the recursive backbone. Trajectory scaling through EMA, cross-run scaling through ensemble distillation, and route scaling each improve performance individually, while combining all three yields the highest AUC.
Model
EMA
KD
Routes
AUC
Recursive backbone
–
–
–
81.59
+ trajectory scaling
✓
–
–
81.65
+ cross-run scaling
–
✓
–
81.68
+ route scaling
–
–
✓
81.61
RECAP
✓
✓
✓
81.76
Table 2: Ablation of estimator-scaling components on the recursive backbone. Trajectory scaling through EMA, cross-run scaling through ensemble distillation, and route scaling each improve performance individually, while combining all three yields the highest AUC.
Model
EMA
KD
AUC
FCN
–
–
81.62
+ trajectory scaling
✓
–
81.67
+ cross-run scaling
–
✓
81.69
+ cross-run + trajectory scaling
✓
✓
81.74
Table 3: Estimator scaling improves the FCN backbone. Trajectory scaling through EMA and cross-run scaling through ensemble distillation each improve AUC individually, while combining both yields the strongest performance.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
AUC (%) ↑
Logloss ↓
EMA decay β
0.9900
81.735
0.43487
0.9950
81.744
0.43482
0.9990
81.760
0.43461
0.9995
81.765
0.43456
0.9999
81.770
0.43450
KD weight λKD
0.00
81.665
0.43535
Appendix
Table 4: Hyperparameter sensitivity of RECAP on Criteo. We vary one method-specific hyperparameter at a time while keeping all other settings fixed. Logloss is the test logloss after validation-fitted affine calibration.
Scaling test-time compute has proven highly effective for language models, yet this opportunity remains largely unexplored for industrial Click-Through Rate (CTR) prediction. CTR models suffer from a fundamental asymmetry: feature combinations well-represented in training yield confident predictions, while sparsely observed ones produce unreliable outputs. Existing training-phase solutions such as adaptive gating learn a fixed selection function subject to the same sparsity, offering no per-instance recourse at deployment.We propose UTTSI (Uncertainty-Triggered Test-Time Selective Inference), a training-free model-agnostic framework that scales inference depth proportionally to per-instance uncertainty. A dual-signal estimator combining model logit confidence with a data-level frequency prior distinguishes epistemic uncertainty from aleatoric ambiguity. Every instance undergoes adaptive feature filtering to remove unreliable embeddings; uncertain instances additionally receive stochastic feature-path explorations whose predictions are aggregated via consistency-weighted ensembling. Confident instances bypass exploration entirely, keeping average overhead at approximately 2.8× base model cost with worst-case latency unchanged.Experiments on four datasets with three backbone architectures demonstrate consistent, statistically significant gains over all training-phase baselines. A seven-day online A/B test further confirms a 5.3% relative CTR gain (p<0.01), establishing selective test-time compute allocation as a practical complement to training-phase advances for CTR prediction.
Generative pre-training via discrete diffusion provides dense reconstruction supervision across all feature fields simultaneously, mitigating representation collapse from data sparsity in CTR prediction. However, all existing generative CTR methods share a fundamental limitation: the reconstruction objective assigns equal training weight to every feature field, ignoring the profound heterogeneity of reconstruction difficulty across high-cardinality ID fields, sparse categorical attributes, numerical values, and behavioral sequences. This causes easy fields to dominate training gradients while the hardest but most informative fields remain chronically underfit, a problem we term the generative difficulty imbalance.We propose HeteGenCTR, which resolves this imbalance through per-field learnable difficulty parameters jointly trained with the denoising network. This unified signal drives two coordinated components without additional hyperparameters: a self-balancing loss that automatically reallocates gradient budget toward harder fields with a provably stable equilibrium, and a difficulty-guided attention mechanism that suppresses the influence of already-converged easy fields while amplifying cross-field information flow toward hard fields. Both components share the same learned signal and remain mutually consistent throughout training. Experiments on five CTR benchmarks and a seven-day online A/B test demonstrate consistent, statistically significant improvements over state-of-the-art baselines, with disproportionate gains for cold-start and long-tail users.
Deep neural models for click-through rate prediction often exhibit a sharp decline in validation performance immediately after the first training epoch despite continued improvement in training loss. This instability restricts effective learning and limits model performance. In this study, we analyze this behavior using large-scale industrial datasets and evaluate practical mitigation strategies. While reducing the learning rate provides only incremental gains, controlling feature sparsity yields substantial improvements. Removing highly sparse features and aggregating infrequent feature values stabilizes training, extends useful learning beyond a single epoch, and improves both offline evaluation metrics and online system performance.
Ergun Biçici, Erkan Çetinyamaç
Intelligent Application Development, Huawei Türkiye R&D Center, Istanbul, Turkey