Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.
Figures & tables
Avazu
Criteo
ML-1M
KDD12
KKBox
Models
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
Logloss ↓
AUC(%) ↑
DNN Covington et al. (2016)
0.3721
79.27
0.4380
81.40
0.3100
90.30
0.1502
80.52
0.4811
85.01
PNN Qu et al. (2016)
0.3712
79.44
0.4378
81.42
0.3070
90.42
0.1504
80.47
0.4793
85.15
Wide & Deep Cheng et al. (2016)
0.3720
79.29
0.4376
81.42
0.3056
90.45
0.1504
80.48
0.4852
85.04
DeepFM Guo et al. (2017)
0.3719
79.30
0.4375
81.43
0.3073
90.51
0.1501
80.60
0.4785
85.31
DCNv1 Wang et al. (2017)
0.3719
79.31
0.4376
81.44
0.3156
90.38
0.1501
80.59
0.4766
85.31
Table 1: Performance comparison of different deep CTR models. Typically, CTR researchers consider an improvement of 0.001 (0.1%) in Logloss and AUC to be practically meaningful Zhu et al. (2021) ; Li et al. (2026) . We conduct a two-tailed t-test over five independent runs to assess the statistical significance between our models and the best baseline (*: p<0.01). Best results are highlighted in bold , second-best are underlined .
Model
EMA
KD
Routes
AUC
Recursive backbone
–
–
–
81.59
+ trajectory scaling
✓
–
–
81.65
+ cross-run scaling
–
✓
–
81.68
+ route scaling
–
–
✓
81.61
RECAP
✓
✓
✓
81.76
Table 2: Ablation of estimator-scaling components on the recursive backbone. Trajectory scaling through EMA, cross-run scaling through ensemble distillation, and route scaling each improve performance individually, while combining all three yields the highest AUC.
Model
EMA
KD
Routes
AUC
Recursive backbone
–
–
–
81.59
+ trajectory scaling
✓
–
–
81.65
+ cross-run scaling
–
✓
–
81.68
+ route scaling
–
–
✓
81.61
RECAP
✓
✓
✓
81.76
Table 2: Ablation of estimator-scaling components on the recursive backbone. Trajectory scaling through EMA, cross-run scaling through ensemble distillation, and route scaling each improve performance individually, while combining all three yields the highest AUC.
Model
EMA
KD
AUC
FCN
–
–
81.62
+ trajectory scaling
✓
–
81.67
+ cross-run scaling
–
✓
81.69
+ cross-run + trajectory scaling
✓
✓
81.74
Table 3: Estimator scaling improves the FCN backbone. Trajectory scaling through EMA and cross-run scaling through ensemble distillation each improve AUC individually, while combining both yields the strongest performance.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
AUC (%) ↑
Logloss ↓
EMA decay β
0.9900
81.735
0.43487
0.9950
81.744
0.43482
0.9990
81.760
0.43461
0.9995
81.765
0.43456
0.9999
81.770
0.43450
KD weight λKD
0.00
81.665
0.43535
Appendix
Table 4: Hyperparameter sensitivity of RECAP on Criteo. We vary one method-specific hyperparameter at a time while keeping all other settings fixed. Logloss is the test logloss after validation-fitted affine calibration.