Assumption-lean logistic regression with missing covariates
Authors: Jyotishka Ray Choudhury, Kabir Aladin Verchand, Richard J. Samworth, Ashwin Pananjady
Organizations: School of Industrial and Systems Engineering, Georgia Institute of Technology · Department of Data Sciences and Operations, University of Southern California · Statistical Laboratory, University of Cambridge · School of Electrical and Computer Engineering, Georgia Institute of Technology
Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex M-estimation problems. These methods and their relatives are suitable for scenarios in which the covariate distribution is known, and more broadly, have enjoyed tremendous success in linear models. But even in basic nonlinear problems such as logistic regression in moderate dimensions, such methods can experience drastic failure modes when the covariate distribution is unknown. Motivated by the need for reliable alternatives, we consider the problem of parameter estimation in logistic regression with missing covariates. Crucially, we operate in the assumption-lean setting where the covariate distribution is unknown (but bounded). We design a stochastic approximation method that is based on Z-estimation with a novel monotone operator, and establish that our algorithm is computationally efficient and achieves provable signal recovery at parametric rates under the hypothesis that covariates are missing completely at random. Our theory sharply characterizes the ℓ22 risk of the estimator in terms of the missingness profile, accommodating heterogeneous observation probabilities. Importantly, it shows that our method always outperforms the de facto ``complete-case'' estimator that ignores observations with any missing data. Even in the setting with homogeneous missingness (in which each covariate is observed independently with probability q), our bounds exhibit intricate and nonstandard dependence on q that can yield significant improvements over using only complete cases. We complement our upper bounds with new information-theoretic lower bounds that show that this intricate dependence on q is fundamental in a minimax sense.
Figures & tables
Figure 1 : Estimation error comparison under MCAR covariate missingness with coordinate-wise observation probability q . The oracle logistic MLE has access to the complete, unmasked covariates, and thus serves as the gold standard. The remaining curves correspond to the complete-case MLE, logistic MLE after zero or Gaussian conditional-mean imputation, the observed-data MLE under a moment-matched Gaussian model, and our proposed SA-based estimator. The curves plot the RMSE ∥θ−θ⋆∥2 over 100 replications against the sample size n . Shaded regions give pointwise 95% confidence intervals across replications.