cs.LGAug 31, 2026

Exact Recovery Thresholds for Weighted Data Selection in Vector-Valued Linear Regression

Authors: Guangjian Zhang

Abstract

We resolve the threshold part of Question 4 of the COLT 2025 open problem "Data Selection for Regression Tasks" of Hanneke, Moran, Shlimovich and Yehudayoff. In vector-valued linear regression with square loss (x,y)(W)=Wxy22\ell_{(x,y)}(W)=|Wx-y|_2^2, where xRdx\in\mathbb{R}^d, yRmy\in\mathbb{R}^m and the learner is the empirical risk minimizer of minimal Frobenius norm, we prove that the minimal budget of weighted examples that recovers the full-data loss on every finite dataset is exactly n(d,m)=(m+1)dn^*(d,m)=(m+1)d. We further determine two more values of the weighted selection profile Fw(d,m,n)F_w(d,m,n): at the near-threshold budget, Fw(d,m,(m+1)d1)=1+1dm2F_w(d,m,(m+1)d-1)=1+\frac{1}{dm^2}, and at the spanning budget, Fw(d,m,d)=d+1F_w(d,m,d)=d+1 for every mm, while Fw(d,m,n)=F_w(d,m,n)=\infty for n<dn<d. For the smallest open intermediate cell (d,m)=(2,2)(d,m)=(2,2) we prove Fw(2,2,3)[13/8,15/8]F_w(2,2,3)\in[13/8,15/8] and Fw(2,2,4)[5/4,3/2]F_w(2,2,4)\in[5/4,3/2], reduce the conjectured exact values 13/813/8 and 5/45/4 to a finite moment problem on the circle with at most seven atoms, and establish strong structural evidence for the conjecture. The upper-bound techniques (a fixed-basis conic compression lemma, a determinant-facet rigidity theorem for maximal certificates, and sharp sparsification lemmas for zero-mean weighted point systems) are of independent interest. As a byproduct we correct an erroneous claim circulating in a recent unrefereed preprint, exhibiting an explicit dataset with m=2m=2 on which no weighted selection of 2d2d points recovers the optimal loss. All results are new only for m2m\ge 2; the scalar case m=1m=1 is due to Hanneke et al.

Explore similar work

CardsList