Semi-Supervised Learning under Spatially Biased Sampling
Organizations: African Institute for Mathematical Sciences (AIMS), 17 KN 16 Ave, Kigali, Rwanda · Mathematical Application & Data Analytics Group, School of Science, Edith Cowan University, 270 Joondalup Drive, Joondalup, WA 6027, Australia · Data & Analytics, Rio Tinto, 152-158 St Georges Terrace, Perth, WA 6000, Australia · Kaplan Business School Pty Ltd, Perth Campus, 1325 Hay St, West Perth, WA, 6005, Australia · Centre for Marine Ecosystems Research, School of Science, Edith Cowan University, 270 Joondalup Drive, Joondalup, WA 6027, Australia
Abstract
Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. This assumption is often violated by biased spatial sampling mechanism, when labels are collected under spatially biased or preferential site selection. We treat this marginal mismatch, spatial autocorrelation, and spatial non-stationarity as three distinct mechanisms, varied independently via a labelled-sampling concentration parameter, a spatial length scale, and a non-stationarity strength parameter, and ask how mismatch degrades SSL, whether the cluster and manifold assumptions survive it, and how the resulting failure can be diagnosed. Using a controlled synthetic framework alongside PovertyMap-WILDS, California housing, socio-economic and US air quality monitoring datasets, we systematically vary the degree of mismatch while accounting for spatial autocorrelation and non-stationarity. Through a series of analyses including a segmented-regression changepoint, we show that in the synthetic generator, SSL performance does not degrade gradually but instead exhibits a threshold-like breakdown between approximately 0.71 and 0.77 once distribution mismatch becomes sufficiently severe. We further demonstrate that spatial non-stationarity contributes to performance loss independently of marginal mismatch and that models become increasingly overconfident outside the regions where labels are available. To support practical deployment, we evaluate several distribution-divergence measures as indicators of reliability and introduce a kernel-weighted local divergence metric that provides a more stable estimate of spatial mismatch than a naïve localised approach. These findings provide empirical evidence and diagnostic tools for better documenting the risk of incorporating unlabelled spatial data into semi-supervised learning workflows.