Abstract
Understanding the fundamental mechanisms of learning is essential for designing systems with strong generalization. Recent studies have shown that increasing the number of training domains, or enlarging the distribution shift among them, improves generalization when each domain contains sufficiently many data samples. However, the conditions under which the data samples can be considered sufficient remain unexplored. In this work, we fill this gap by establishing criteria for per-domain sample requirements based on the presented learning bounds. These criteria not only reveal an inverse linear scaling law between the number of training domains and the number of samples required per domain, but also explain the fundamental rationale behind the assumption of data sufficiency, thereby providing theoretical guidance for assessing the adequacy of existing datasets and constructing datasets. This differs from classical learning theory, as the number of samples required is highly dependent on the number of training domains. Additionally, we prove the close relationship between in-domain learning and out-of-domain generalization through the presented generalization bounds, and lastly discuss some key arguments.
Explore similar work
Jul 17, 2026cs.LG
We study hierarchical domain generalization as a problem of extrapolation from finite observed regions to an entire instance space, replacing i.i.d. sampling with arbitrary domain hierarchies. We show that the central obstruction is not only the complexity of the hypothesis class, but the train/test domain partition through which evidence is revealed. In particular, no matter how small the class or how large the training size, some partition makes generalization fail for some target. These results suggest that modern generalization theory must treat domain structure as a first-class object.
Chenxiao Yang, Zhiyuan Li, Shai Ben-David +1
Toyota Technological Institute at Chicago · University of Waterloo
Sep 30, 2026cs.LG
The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation and generalization. Structural Risk Minimization (SRM) has long been regarded as a principled solution under the classical i.i.d. assumption. However, domain generalization (DG) violates this assumption, leaving the theoretical role of SRM in DG largely unexplored. To bridge this gap, we establish the first theoretical guarantees for SRM in DG under mild assumptions. Specifically, based on the concept of stability, we derive learning consistency and generalization error bounds and prove that these bounds become tight when the hypotheses satisfy the stability condition. Building upon this, under a specific hypothesis space assumption, we establish stability, learning, and generalization bounds for SRM. We further discuss the applicability of these bounds to deep learning. This work establishes theoretical foundations for SRM under distribution shifts and sheds light on the design of robust DG algorithms in small-sample scenarios.
Hong Zheng
This work was completed during my Ph.D. period and was the result of a flash of inspiration. I decided to submit it to arXiv and leave its value to the readers.
Jun 6, 2026cs.LG
Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understanding of these model loss behaviors remains absent. In this work, we propose a unified framework to explain the underlying mechanics of data mixing. Our approach extends theoretical perspectives originally developed for standard neural scaling laws (e.g., Kaplan and Chinchilla) to the multi-domain setting. Based on the distributional assumption that domains overlap on fundamental skills while diverging on specialized skills, we identify two key factors that govern the domain losses of models trained on different data mixtures: \textit{Capacity Competition}, where the allocation of finite model capacity couples domain losses globally, and \textit{Noise Reduction}, where optimal weights shift toward harder-to-learn domains to minimize overall noise. Empirical evaluations show that our framework outperforms existing baselines by fitting the loss landscape with a lower Mean Relative Error and identifying higher-performing training mixtures. Most importantly, our model successfully extrapolates across scales, predicting highly effective mixtures for large, unseen scales using parameters fitted on smaller ones. In addition, our model achieves these results using significantly fewer parameters compared to previous empirical laws. Our code is available at https://github.com/meiqwq/Explaining-Data-Mixing-Scaling-Laws.
Rui Dai, Shuran Zheng
Beijing Institute of Technology · IIIS, Tsinghua University