When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
Authors: Martin Hofmann, Patrick Mäder
Organizations: Data-intensive Systems and Visualization Group (dAI.SY), Technische Universität Ilmenau, Max-Planck-Ring 14, 98693 Ilmenau, Thuringia, Germany · German Centre for Integrative Biodiversity Research (iDiv) Halle–Jena–Leipzig, Deutscher Platz 5e, 04103 Leipzig, Saxony, Germany · Faculty of Biological Sciences, Friedrich Schiller University, Fürstengraben 1, 07745 Jena, Thuringia, Germany
Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.
Figures & tables
Figure 1: Design of the main study. A network is trained on 12 tasks and then on a common anchor task. At each of four checkpoints, a copy receives 100 updates on a new task and its loss reduction is recorded. History models see one event per 100 updates of the history together with the diagnostics of the current checkpoint. The baseline sees the diagnostics only.
Criterion, in the fixed order
Requirement
Result
Outcome
Known answer
exact within 10−12
reproduced
met
Repeatability, minimum over classes
ICC(1,1) ≥0.80
0.781 [0.673, 0.983]
not met
Gain of the compact state over B1
≥10% , lower bound >0
− 21.4% [ − 91.9, 8.1]
not met
Transfer to a held-out regime
mean gain ≥5% , each regime >0
mean − 80.3%
not met
Against the flexible baseline
RMSE difference, lower bound >0
0.537 [ − 0.412, 0.989]
not met
Dimension of the response map
stable under resampling
three components, stable
met
Table 1: Criteria of the main study in the fixed order. Below the line: the positive control in the second stage and the declared third repeat. Intervals are 90% bootstrap intervals over histories.
Figure 2: Positive control in the main study. a, b: original design on six histories. Lines are medians of the response distance to the unchanged checkpoint (signal), of the distance between two unchanged measurements (noise floor) and of their difference z , for averaging windows ending at each step; bands are 90% bootstrap intervals. c: second stage with stronger doses on 12 histories, window ending at step 85. Points are histories, bars are medians with 90% intervals. The dotted line is the threshold fixed before the first run.
Figure 3: Prediction of the response to further training in the main study. a: error on the six test histories; the dotted line marks the current-state baseline. b: gain over the baseline with 90% bootstrap intervals over histories; the dashed line is the gain required in advance. c: gain of the compact state when a whole regime is held out; the dashed line is the required mean gain.
1,000 runs for training the forecaster
200 runs
Predictor
epoch 12
24
48
12
24
48
Current value, unfitted
0.680
0.496
0.319
0.680
0.496
0.319
Snapshot model
0.545
0.434
0.378
0.629
0.551
0.488
Compact recurrent state
0.474
0.392
0.316
0.483
0.398
0.368
GRU over history
0.483
0.398
0.345
0.509
0.412
0.370
Transformer over history
0.481
0.401
0.321
0.533
0.434
0.385
Table 2: Forecasting screen: mean absolute error of the forecast of the final log test error, with one task family held out (average over 13 families). Bold marks the lowest error in a column.
Figure 4: Forecasting screen with 1,000 runs for training the forecasters. a: forecast error with one task family held out. b: improvement of the three history models over the snapshot model (left) and over the current validation error (right), with 90% bootstrap intervals over families. The dashed line is the threshold set for epoch 48 against the snapshot model.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Repeatability of the response in the main study, with 90% bootstrap intervals over histories. The single conflict measurement missed the criterion. The value for the mean of two repeats is a projection; the value for three repeats was measured afterwards.
Finding
Obtained
Units
Main study: no gain of the compact state or the history GRU
fixed in advance
6 test histories
Main study: no transfer to a held-out regime
fixed in advance
3 regimes
Main study, stage 1: conflict response below the reliability criterion
fixed in advance
24 checkpoints
Main study, stage 2: dose response of the positive control
fixed before new data, after the first result
12 histories
Main study: reliability of the three-repeat mean
extension declared before execution
24 checkpoints
Main study: re-initialisation control, paired
technical control frozen before production
6 pairs
Appendix
Table 3: Status of the findings.
Predictor
epoch 6
epoch 12
epoch 24
Current accuracy, unfitted
0.2865
0.2201
0.1363
Current accuracy, linear calibration
0.0844
0.0619
0.0437
Current accuracy and architecture, quadratic †
0.0669
0.0492
0.0374
Latest multichannel snapshot
0.0869
0.0681
0.0717
Order-free summary of the prefix
0.0791
0.0884
0.0687
Reservoir, best of ten arms †
0.0684
0.0661
0.0496
Appendix
Table 4: CIFAR-100 screen: mean absolute error of the forecast of final test accuracy (as a fraction) with a block of tasks held out. † : selected or added after the first analysis. For the snapshot, summary and reservoir rows, the better of the available readouts on the test folds is shown for each column.
Figure 6: Left: CIFAR-100 screen, forecast error with a block of tasks held out; the unfitted current accuracy (error 0.287 to 0.136) is off the scale. Right: class-incremental screen, mean forecast error over five targets in percentage points with a class order held out.
Predictor
u=0
1
2
5
10
20
Current values, calibrated
0.920
0.916
0.915
0.891
0.867
0.658
Snapshot model
0.851
0.846
0.878
0.813
0.794
0.683
Order-free summary of the prefix
2.372
1.868
1.796
1.249
1.846
1.308
Echo state network
1.821
1.727
1.722
1.714
1.776
1.661
Compact recurrent state
0.985
0.936
1.006
1.010
0.966
0.925
GRU over history
0.856
0.878
0.878
0.865
0.832
0.796
Appendix
Table 5: Class-incremental screen: mean absolute forecast error in percentage points, averaged over five targets, after u updates on the new class, for the input set with current errors and training telemetry. Bold marks the lowest error in a column.
The linear recurrent neural network (LRNN) is a simple model for studying how much memory a network builds up as it trains. For uncorrelated inputs, earlier work found that training itself settles the network between keeping the past and reacting only to the present. Real sequences are correlated, and we solve the learning dynamics exactly for correlated inputs. In the solution, keeping the past carries a cost. The whole effect of correlation lands on that cost. This cost reduces to the earlier one when inputs are uncorrelated and grows once they are positively correlated. Three findings follow. (1) Correlation reshapes the course of learning, not only its end. Memory builds, overshoots, and is partly removed, and the settled network keeps less of the past. (2) Memory switches off at a threshold set by one number, how much each input resembles the one just before it. Neither sequence length nor longer-range correlation moves this threshold. Memory is worth keeping only when the task needs the previous input more than the current input already supplies it through correlation with the past. (3) The best network changes too. Zero error demands a feedthrough, a path that passes the current input straight to the network's output and remembers nothing, and training builds it unprompted when given one spare hidden dimension. Our work turns one property of the input into a prediction of whether a network learns memory and explains why correlated data turns recurrent networks into change detectors.
Arnol Manuel Fokam, Fasseu Sieyondji Akpevwoghene, Edem Fiifi Dawson
Independent Researcher, United Kingdom · Department of Electrical and Electronic Engineering, University of Buea, Cameroon · minoHealth AI Labs, Ghana
A temporally drifting data stream may pass through discrete regimes rather than changing continuously. We ask whether such regimes are recoverable from the weights of models trained on the stream, using a hidden Markov model (HMM) fit to the chronologically ordered trajectory of those weights. We study this question in two domains known to drift over time: multimodal misinformation detection, using the Fakeddit dataset; and sentiment analysis, using the Yelp dataset. We train classifiers on consecutive temporal windows and fit an HMM to the trajectory of their aligned weights, recovering latent states that partition each timeline into coherent phases. On both datasets, classifiers generalize better to data from windows sharing the state of their training window than to windows across state boundaries. This within-state transfer advantage survives a control for temporal proximity and modestly exceeds the advantage recovered by a naive partition into contiguous states of equal size. Although the states are estimated solely from model weights, they correlate more strongly with shifts in the data's class distribution than with the weight-space geometry used to estimate them. After class divergence and lag are residualized out, the within-state advantage exceeds its permutation null on both tasks, indicating that the states recover structure relevant to transfer beyond the data distribution. Every effect replicates on both tasks but is attenuated on Yelp, whose label distribution is more temporally stable.
Predicting observed dynamics does not establish recovery of the underlying physical mechanism. Can machine learning retrace the hidden-state reasoning behind the Hodgkin-Huxley (HH) model? We train structured latent models on simulated current and voltage, withholding gate identities and trajectories from training and model selection. We then test response prediction, state recovery, protocol transfer, and agreement with HH dynamics. Prediction error and its cross-seed spread both drop sharply at three latent dimensions under the tested protocols, while gate recovery under new protocols improves through five to six coordinates. State recovery depends on which observations the chart uses. Observed voltage improves current-clamp decoding relative to freely predicted voltage. Under voltage clamp, adding latent state to command voltage raises m-state R2 from 0.976 to above 0.99, yet the transported field disagrees with HH on identical smooth samples. Known invertible HH coordinates achieve high fast-m field agreement under the same audit procedure. An exact HH identity decomposes the discrepancy into time-scale-weighted state error and a residual in the transported field; these terms can cancel or reinforce. These findings concern the tested models and charts. They support evaluating state and dynamics recovery separately, including chart inputs and transported-field agreement across interventions.
Peiyu Zang, Jiayi Hao, Yongqiang Cai
School of Mathematical Sciences, Beijing Normal University