cs.LGSep 29, 2026

Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases

Authors: Vikram Kher, Jane H. Lee, Anay Mehrotra, Manolis Zampetakis

Organizations: Yale University · Stanford University

Abstract

When can we learn from biased samples? We study regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes themselves, a ubiquitous challenge spanning clinical trials with patient dropout, labor markets with self-selection, and auctions with strategic entry. Ignoring such selection yields systematically biased conclusions with real-world consequences. This challenge has a long history in econometrics and statistics, starting with Heckman's seminal two-stage model and followed by numerous generalizations. While these works provide various sufficient conditions for identification, a complete characterization of when such regression is possible has remained elusive. In this work, we provide a characterization for when regression is possible in the presence of sample selection bias. Our results establish the minimal assumptions required on the functional forms of selection processes under which regression remains possible, which are particularly relevant in modern settings where selection mechanisms are increasingly complex and opaque. As a corollary of our characterization, we show that there are settings where the regression function can be identified even when the selection filter itself cannot. This observation already goes beyond the ``estimate selection filter, then debias regression'' paradigm that is followed by virtually all existing approaches. Under natural strengthenings of our identification conditions, we also establish finite-sample estimation guarantees with explicit convergence rates and provide oracle-efficient algorithms. This yields the first general-purpose estimation method for this broad class of selection problems. Finally, we explore the implications of our results for several well-studied econometric settings with complex selection mechanisms such as auctions with entry costs and labor markets.

Explore similar work

May 13, 2026stat.ME

Towards a holistic understanding of Selection Bias for Causal Effect Identification

Selection bias is pervasive in observational studies. For example, large scale biobanks data can exhibit ``healthy volunteer bias'' when respondents are healthier and of higher socio-economic status than the population they are meant to represent. Recovering causal effects from such sub-population is an important problem in causal inference, as estimating average treatment effects (ATE) from selected populations can result in a severely biased estimate of the ATE from the whole population. In this paper, we investigate the identifiability of the ATE under selection bias. We provide necessary and sufficient conditions for ATE identifiability, leveraging weak assumptions on probability classes to characterize propensity score and selection probability. Compared to previous works, our results extend existing graphical identifiability criteria and offer a more comprehensive understanding of causal effect identification with strictly weaker conditions in the presence of selection bias.
Jun 12, 2026econ.EM

Identification and Inference for Algorithmic Frontiers with Selective Labels

This paper provides identification results to characterize a fairness-accuracy (FA) frontier, and statistical inference tools to test hypotheses and build a confidence set for the FA-frontier, when outcomes are observed only for selected individuals. When the selection process is unrestricted but loss is measured in specific ways, we provide a characterization of the sharp identification region of the FA-frontier. Under an assumption of unconfoundedness conditional on observables (and unrestricted loss functions), we obtain point identification and propose a debiased machine learning estimator, derive its asymptotic distribution, and show how this can be used to carry out inference for the FA-frontier. In work in progress, we extend the partial identification results to a broader class of loss functions.
Date pendingstat.ML

Can SGD Select Good Fishermen? Local Convergence under Self-Selection Biases

We revisit the problem of estimating kk linear regressors with self-selection bias in dd dimensions with the maximum selection criterion, as introduced by Cherapanamjeri, Daskalakis, Ilyas, and Zampetakis [CDIZ23, STOC'23]. Our main result is a poly(d,k,1/ε)+(klog⁡k)O(k)\mathrm{poly}(d, k, 1/\varepsilon) + (k \log k)^{O(k)} time algorithm for this problem that improves upon the running time of the algorithms by Cherapanamjeri, Daskalakis, Ilyas, and Zampetakis [CDIZ23] and Gaitonde and Mossel [GM24, arXiv]. We achieve this by providing the first local convergence algorithm for self-selection, thus resolving one of the main open questions of Cherapanamjeri, Daskalakis, Ilyas, and Zampetakis [CDIZ23]. To obtain this algorithm, we reduce self-selection to a seemingly unrelated statistical problem called estimation under coarsening [FKKT21, COLT'21]. Coarsening occurs when one does not observe the exact value of the sample but only some set (from a partition of the sample space) containing the exact value. Inference from coarse samples arises in various real-world applications, including rounding by humans and algorithms, limited precision of instruments, and lag in multi-agent systems. The coarse estimation problem arising in our reduction is induced by a non-convex partition, whereas previous works on coarsening exclusively studied convex partitions. The resulting estimation algorithm relies on the geometry of the self-selection problem to bypass non-convexity. This geometric approach, in turn, enables us to overcome the limitations of previous analytic approaches and could have applications for designing efficient algorithms for other latent-variable problems.