End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.
Figures & tables
Figure 1: Greedy layer-wise vs. end-to-end training. Left: Demonstration of greedy training. As an example, we show what happens while training Layer 3 (the same procedure is applied to the other layers). In this example, Layers 1–2 are frozen, only Layer 3 is trained, and the gradient stops at Layer 2. Layer 4 is not yet added to the network; it will be trained once Layer 3 is trained. Right: End-to-end training. All layers are trained jointly, and the gradient reaches every layer.
Figure 2: Increased network width narrows the performance gap between greedy layer-wise and end-to-end training. Conv4 and Conv8 networks trained with Barlow Twins loss on the CIFAR-10 dataset. (a,b) Terminal and best-epoch test k NN accuracy across widths for Conv4 (a) and Conv8 (b). Solid lines indicate terminal accuracy; dotted lines indicate the highest recorded accuracy. (c) Example test k NN accuracy training trajectory for Conv4 at 32× width. (d) Change in terminal accuracy relative to each method’s 1× baseline.
Figure 3: Lower Barlow Twins loss does not necessarily imply higher categorization accuracy. (a) Terminal test k NN accuracy versus final training loss across network widths. Each point represents one width; connecting lines follow increasing width. The loss axis decreases from left to right. (b) Training loss curves for models with 16× width for both Conv4 and Conv8 architectures with either greedy layer-wise training or end-to-end (E2E) training. The loss axis (y) is logarithmic. Dotted vertical lines indicate layer additions for greedy training.
Figure 4: Representational geometry across network widths and training methods. Signal–signal factorization (SSF) and signal–noise factorization (SNF) for Conv4 networks trained to minimize Barlow Twins loss on CIFAR-10. (a,b) Final encoder-layer SSF and SNF across widths. (c,d) SSF and SNF over 1,000 epochs for Conv4 32× width.
Figure 5: Width-dependent performance trends extend to another self-supervised loss function and dataset. Terminal and best-epoch test k NN accuracy for Conv4 networks trained end-to-end or greedily layer-wise on a single seed. (a) k NN categorization accuracy for Conv4 models trained to minimize SimCLR loss on CIFAR-10 across widths from 0.25× to 32× . (b) k NN categorization accuracy for Conv4 models trained to minimize Barlow Twins loss on the CIFAR-100 dataset across widths from 0.25× to 32× . Solid lines indicate terminal accuracy; dotted lines indicate the highest recorded accuracy at full encoder depth.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Recipe
1×
8×
Mean
Adam 10−3 , constant
67.12
75.36
71.24
Adam 3×10−3 , cosine
66.52
73.88
70.20
Adam 10−3 , cosine
65.82
73.82
69.82
Adam 3×10−4 , cosine
62.20
71.74
66.97
Adam 10−4 , cosine
57.78
68.88
63.33
LARS, warmup + cosine
58.06
66.90
62.48
Appendix
Table 1: Initial end-to-end optimizer screen with terminal validation k NN accuracy. Reported values are top-1 classification accuracy on CIFAR-10 using a k NN probe.
Width
Adam 10−3 const (end-to-end)
Adam 10−3 cosine (end-to-end)
Adam 10−3 const (greedy)
0.25×
55.37
55.52
50.40
0.5×
61.70
61.62
56.95
1×
67.69
66.58
63.60
2×
72.11
71.36
67.70
4×
75.85
75.43
71.98
8×
77.32
77.88
75.37
Appendix
Table 2: Additional learning-rate schedule test for end-to-end learning with Adam 10−3 cosine and constant. Results of greedy layerwise training are included for reference. Reported values are terminal validation k NN probe accuracy on CIFAR-10.
q
λ
Candidate source
1×
8×
Mean
256
0.0051
Fixed
64.78
72.62
68.70
256
0.0200
Ghosh/FastSSL
65.40
74.92
70.16
256
0.1632
Scaled
66.16
75.98
71.07
1024
0.0051
Fixed
66.48
74.64
70.56
1024
0.0020
Ghosh/FastSSL
66.14
74.14
70.14
1024
0.0408
Scaled
66.78
74.90
70.84
Appendix
Table 3: Barlow Twins projector/lambda screen for models trained end-to-end. Reported values are terminal validation k NN probe accuracy on CIFAR-10.
Width
q
λ
Seed 0
Seed 1
Mean ± SD
1×
256
0.1632
66.16
65.90
66.03±0.18
1×
1024
0.0408
66.78
65.26
66.02±1.07
8×
256
0.1632
75.98
75.78
75.88±0.14
8×
1024
0.0408
74.90
75.04
74.97±0.10
Appendix
Table 4: Additional-seed confirmation of the two leading projector/lambda configurations for end-to-end training. Reported values are terminal validation k NN probe accuracy on CIFAR-10.
Optimizer
Schedule
1×
8×
Mean
Adam 10−3
Constant
64.98
74.42
69.70
Adam 3×10−3
Constant
64.52
74.10
69.31
Adam 10−3
Cosine to 10−6
63.06
72.04
67.55
LARS
warmup + cosine
52.78
63.02
57.90
Appendix
Table 5: Barlow Twins optimizer/schedule screen. Entries are terminal validation k NN probe accuracy (%) on CIFAR-10 after end-to-end training.
End-to-end
Greedy layer-wise
Width
8% crop
20% crop
8% crop
20% crop
0.25×
58.75
58.11
53.73
55.74
0.5×
63.80
64.00
61.00
60.90
1×
68.40
69.15
65.88
66.87
2×
72.41
72.84
69.37
70.59
4×
74.28
75.23
72.25
73.52
Appendix
Table 6: Effect of minimum crop-area fraction on terminal test k NN probe accuracy (%) for Conv4 trained end-to-end to minimize SimCLR loss on CIFAR-10. Greedy layer-wise is included as reference.
Hidden dimension
1×
8×
Mean
128
63.28
70.16
66.72
256
64.02
70.68
67.35
512
64.40
71.10
67.75
1024
64.36
70.98
67.67
2048
64.88
71.80
68.34
Appendix
Table 7: Projector hidden-dimension screen for models trained end-to-end with SimCLR loss. Reported values are terminal validation k NN probe accuracy on CIFAR-10.