Scalable Cox Regression via Grouped Risk Sets and Sharper LogSumExp Rates
Organizations: HSE University
Abstract
Motivated by the computational challenges of large-scale Cox regression, we study stochastic minimization of LogSumExp objectives over large sets. Mini-batch normalizer estimates generally yield biased gradients. We instead use a softplus surrogate that introduces one auxiliary scalar per normalizer and admits unbiased single-sample gradients. For smooth convex LogSumExp objectives, we prove an averaged objective bound, improving the previous analysis. With a strongly convex regularizer on the original variable, we also obtain a last-iterate squared-error rate of without strong convexity in the auxiliary variables. For Cox regression, the normalizers are defined over nested risk sets. We exploit this structure by grouping neighboring failures and sharing one auxiliary variable per group. The resulting compressed objective admits uniform score and curvature bounds that control the errors from grouping and softplus approximation. Together with the general optimization result, these bounds give a mean-square rate of , up to logarithmic factors, relative to the full Cox solution. The compressed estimator also matches the full estimator's asymptotic distribution. Experiments on synthetic and real survival datasets with slowly decreasing risk sets show a favorable performance relative to stochastic baselines.
Figures & tables
| Dataset | Tuning cap | Final cap | ||||
|---|---|---|---|---|---|---|
| SUPPORT2 | 8,873 | 22 | 3,621 | 61 | 3 | 4.5 |
| NWTCO | 4,028 | 11 | 342 | 15 | 3 | 4.5 |
| Correlated 100k | 100,000 | 10 | 6,005 | 133 | 10 | 15 |
| Correlated 1M | 1,000,000 | 20 | 60,136 | 178 | 100 | 150 |
| Independent 100k | 100,000 | 20 | 41,921 | 166 | 10 | 15 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Candidates |
|---|---|
| Ours | |
| Batch LSE | : ; : |
| Minibatch Cox | |
| BigSurv | Stratum size 20; |
| Cox-CC | ; |
| Method | SUPPORT2 | NWTCO | Correlated 100k | Correlated 1M | Independent 100k |
|---|---|---|---|---|---|
| Ours | |||||
| Batch LSE | |||||
| Minibatch Cox | |||||
| BigSurv | |||||
| Cox-CC |
| Dataset | Ours | Batch LSE | Minibatch Cox | BigSurv | Cox-CC |
|---|---|---|---|---|---|
| SUPPORT2 | |||||
| NWTCO | |||||
| Correlated 100k | |||||
| Correlated 1M | |||||
| Independent 100k |
| Dataset | Method | Cost (M) | Cost (M) | ||
|---|---|---|---|---|---|
| SUPPORT2 | Ours | 10/10 | 0.3115 | 0/10 | – |
| SUPPORT2 | Batch LSE | 1/10 | 4.4980 | 0/10 | – |
| SUPPORT2 | Minibatch Cox | 0/10 | – | 0/10 | – |
| SUPPORT2 | BigSurv | 10/10 | 0.0654 | 0/10 | – |
| SUPPORT2 | Cox-CC | 3/10 | 4.4999 | 0/10 | – |
| NWTCO | Ours | 10/10 | 0.3108 | 0/10 | – |
| Dataset | Method | Test Cox loss | Test C-index |
|---|---|---|---|
| SUPPORT2 | Ours | ||
| SUPPORT2 | Batch LSE | ||
| SUPPORT2 | Minibatch Cox | ||
| SUPPORT2 | BigSurv | ||
| SUPPORT2 | Cox-CC | ||
| NWTCO | Ours |
| Dataset | Method | Validation Cox loss | Validation C-index |
|---|---|---|---|
| SUPPORT2 | Ours | ||
| SUPPORT2 | Batch LSE | ||
| SUPPORT2 | Minibatch Cox | ||
| SUPPORT2 | BigSurv | ||
| SUPPORT2 | Cox-CC | ||
| NWTCO | Ours |
| Dataset | slope | slope | |
|---|---|---|---|
| SUPPORT2 | 61 | 1.783 | 1.000 |
| NWTCO | 15 | 1.601 | 0.997 |
| Correlated 100k | 133 | 1.372 | 0.999 |
| Correlated 1M | 178 | 1.716 | 0.999 |
| Independent 100k | 166 | 1.866 | 0.999 |