When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
Authors: Martin Hofmann, Patrick Mäder
Organizations: Data-intensive Systems and Visualization Group (dAI.SY), Technische Universität Ilmenau, Max-Planck-Ring 14, 98693 Ilmenau, Thuringia, Germany · German Centre for Integrative Biodiversity Research (iDiv) Halle–Jena–Leipzig, Deutscher Platz 5e, 04103 Leipzig, Saxony, Germany · Faculty of Biological Sciences, Friedrich Schiller University, Fürstengraben 1, 07745 Jena, Thuringia, Germany
Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.
Figures & tables
Figure 1: Design of the main study. A network is trained on 12 tasks and then on a common anchor task. At each of four checkpoints, a copy receives 100 updates on a new task and its loss reduction is recorded. History models see one event per 100 updates of the history together with the diagnostics of the current checkpoint. The baseline sees the diagnostics only.
Criterion, in the fixed order
Requirement
Result
Outcome
Known answer
exact within 10−12
reproduced
met
Repeatability, minimum over classes
ICC(1,1) ≥0.80
0.781 [0.673, 0.983]
not met
Gain of the compact state over B1
≥10% , lower bound >0
− 21.4% [ − 91.9, 8.1]
not met
Transfer to a held-out regime
mean gain ≥5% , each regime >0
mean − 80.3%
not met
Against the flexible baseline
RMSE difference, lower bound >0
0.537 [ − 0.412, 0.989]
not met
Dimension of the response map
stable under resampling
three components, stable
met
Table 1: Criteria of the main study in the fixed order. Below the line: the positive control in the second stage and the declared third repeat. Intervals are 90% bootstrap intervals over histories.
Figure 2: Positive control in the main study. a, b: original design on six histories. Lines are medians of the response distance to the unchanged checkpoint (signal), of the distance between two unchanged measurements (noise floor) and of their difference z , for averaging windows ending at each step; bands are 90% bootstrap intervals. c: second stage with stronger doses on 12 histories, window ending at step 85. Points are histories, bars are medians with 90% intervals. The dotted line is the threshold fixed before the first run.
Figure 3: Prediction of the response to further training in the main study. a: error on the six test histories; the dotted line marks the current-state baseline. b: gain over the baseline with 90% bootstrap intervals over histories; the dashed line is the gain required in advance. c: gain of the compact state when a whole regime is held out; the dashed line is the required mean gain.
1,000 runs for training the forecaster
200 runs
Predictor
epoch 12
24
48
12
24
48
Current value, unfitted
0.680
0.496
0.319
0.680
0.496
0.319
Snapshot model
0.545
0.434
0.378
0.629
0.551
0.488
Compact recurrent state
0.474
0.392
0.316
0.483
0.398
0.368
GRU over history
0.483
0.398
0.345
0.509
0.412
0.370
Transformer over history
0.481
0.401
0.321
0.533
0.434
0.385
Table 2: Forecasting screen: mean absolute error of the forecast of the final log test error, with one task family held out (average over 13 families). Bold marks the lowest error in a column.
Figure 4: Forecasting screen with 1,000 runs for training the forecasters. a: forecast error with one task family held out. b: improvement of the three history models over the snapshot model (left) and over the current validation error (right), with 90% bootstrap intervals over families. The dashed line is the threshold set for epoch 48 against the snapshot model.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Repeatability of the response in the main study, with 90% bootstrap intervals over histories. The single conflict measurement missed the criterion. The value for the mean of two repeats is a projection; the value for three repeats was measured afterwards.
Finding
Obtained
Units
Main study: no gain of the compact state or the history GRU
fixed in advance
6 test histories
Main study: no transfer to a held-out regime
fixed in advance
3 regimes
Main study, stage 1: conflict response below the reliability criterion
fixed in advance
24 checkpoints
Main study, stage 2: dose response of the positive control
fixed before new data, after the first result
12 histories
Main study: reliability of the three-repeat mean
extension declared before execution
24 checkpoints
Main study: re-initialisation control, paired
technical control frozen before production
6 pairs
Appendix
Table 3: Status of the findings.
Predictor
epoch 6
epoch 12
epoch 24
Current accuracy, unfitted
0.2865
0.2201
0.1363
Current accuracy, linear calibration
0.0844
0.0619
0.0437
Current accuracy and architecture, quadratic †
0.0669
0.0492
0.0374
Latest multichannel snapshot
0.0869
0.0681
0.0717
Order-free summary of the prefix
0.0791
0.0884
0.0687
Reservoir, best of ten arms †
0.0684
0.0661
0.0496
Appendix
Table 4: CIFAR-100 screen: mean absolute error of the forecast of final test accuracy (as a fraction) with a block of tasks held out. † : selected or added after the first analysis. For the snapshot, summary and reservoir rows, the better of the available readouts on the test folds is shown for each column.
Figure 6: Left: CIFAR-100 screen, forecast error with a block of tasks held out; the unfitted current accuracy (error 0.287 to 0.136) is off the scale. Right: class-incremental screen, mean forecast error over five targets in percentage points with a class order held out.
Predictor
u=0
1
2
5
10
20
Current values, calibrated
0.920
0.916
0.915
0.891
0.867
0.658
Snapshot model
0.851
0.846
0.878
0.813
0.794
0.683
Order-free summary of the prefix
2.372
1.868
1.796
1.249
1.846
1.308
Echo state network
1.821
1.727
1.722
1.714
1.776
1.661
Compact recurrent state
0.985
0.936
1.006
1.010
0.966
0.925
GRU over history
0.856
0.878
0.878
0.865
0.832
0.796
Appendix
Table 5: Class-incremental screen: mean absolute forecast error in percentage points, averaged over five targets, after u updates on the new class, for the input set with current errors and training telemetry. Bold marks the lowest error in a column.