cs.LGOct 1, 2026

Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

Authors: Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran

Organizations: School of Computer Science, The University of Sydney, Sydney, Australia · School of Computer Science and Engineering, UNSW Sydney, Sydney, Australia

Abstract

Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from p=.3p=.3 to .5.5 to .7.7, while preserving the relative mixture 1:2:3:41{:}2{:}3{:}4 among the four shared feature levels, increases the final mean accuracy gap from .185.185 to .331.331 to .527.527 across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Misalignment of Low-Loss Regions Causes Grokking

    Sep 30, 2026Yongding Tian, Zaid Al-Ars, Maksim Kitsak +1GrokkingOverfitting

  2. Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

    May 12, 2026Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer +2Model Fine-TuningRetention

  3. A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to Adapt

    Jun 14, 2026Tomoya WakayamaStable Test-Time AdaptationTest-Time Training