stat.MLOct 7, 2026

Dataset Pruning from First Principles: A Label-Free Linear Programming Approach

Authors: Rodrigo Schuller, Francisco Ganacim

Organizations: Instituto de Matem´atica Pura e Aplicada (IMPA) Rio de Janeiro, Brazil

Abstract

Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather than imposing this assumption, we derive geometric selection criteria by reformulating unbiased subset selection as a variance minimization problem. Unbiasedness ensures that unweighted subset averages recover full-dataset averages in expectation, including losses and gradients at fixed model parameters. Specifically, we characterize a family of unbiased subset selection algorithms as a high-dimensional polytope. In this context, minimizing the expected sampling variance is a linear objective. Differences in sampling variance, averaged over rigid motions, admit closed-form pairwise expressions. Because the polytope has high dimension, directly applying standard linear programming is impractical. We instead use these expressions to construct an efficient vertex walk that optimizes an approximation of the variance objective while preserving unbiasedness, yielding a method that requires neither labels nor model training during selection. Across CIFAR-10, MNIST, and CelebA benchmarks, our method matches or exceeds uniform sampling in mean test accuracy at every evaluated budget and outperforms competing geometric methods in several settings, particularly at small selection budgets. Beyond dataset pruning, the same framework reduces stochastic-gradient variance by increasing diversity within mini-batches while keeping the batch size unchanged.

Figures & tables

Explore similar work

CardsList
  1. Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

    Jun 11, 2026Dongyue Wu, Zilin Guo, Xiaoyu Li +4Structured PruningGraph Representation Learning

  2. Data Pruning: Redundant, Problematic, and Interdependent Samples

    Jun 20, 2026Leon Freese, Marthinus W. TheunissenNoisy LabelsData Quality

  3. Label-Efficient Dataset Pruning via Semi-Supervised Pseudo-Labeling

    May 22, 2026Yeseul Cho, Baekrok Shin, Changmin Kang +1High Confidence Pseudo-LabelsSemi-Supervised Learning