Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the projector. We find that this mismatch does not necessarily prevent dimensional collapse in the backbone, which can retain low effective rank and potentially limit downstream transfer. To address this, we introduce SACReg, a spectral anti-collapse regularizer motivated by an analysis of λ-balance, which captures the relative scale of weight matrices across layers. In a two-layer linear network, we show that (i) λ-balance prevents collapse, and (ii) our regularizer applied to the backbone induces λ-balance. In the nonlinear case, this regularizer leads to anti-collapse as well and, in realistic architectures on ImageNet100, it empirically increases the representations' ranks. We apply SACReg to JEPA and propose λ-JEPA, which improves over LeJEPA and VISReg on ImageNet-1k classification and in average linear-probe transfer performance across eight downstream image datasets. On video self-supervised learning, λ-JEPA improves over LeVJEPA and V-JEPA 2 on the Something-Something-v2 and Kinetics-400 benchmarks. Code is available at https://github.com/berkerdemirel/lambda-jepa.
Figures & tables
Figure 1: Backbone and loss space geometry for explicitly regularized JE-SSL methods. The projection head changes both pairwise similarities and RankMe.
Method
IN-1k
SSv2
K400
ViT-S/16, 240 epochs
LeVJEPA
39.4
—
—
V-JEPA 2
38.7
—
—
λ -JEPA
46.8
37.3
40.0
ViT-B/16, 240 epochs
VideoMAEv2
47.1
—
—
Table 2: Video SSL on IN-1k, SSv2, and K400. Longer training LeVJEPA and λ -JEPA use 1085 epochs.
Figure 2: Effect of backbone SACReg across SSL objectives on ImageNet-100. Top: backbone statistics and downstream evaluation. Bottom: class-center separability versus view sensitivity along class-prediction directions. Open and green markers denote baseline and +SACReg, respectively; arrows show their change. Detailed results are in Appendix C.6 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Block Gram matrix Q=(W1⊤W1MM⊤W2W2⊤) , M=W2W1, for a two-layer linear network. (A) Training with symmetric L2 weight decay ( λreg=0.1 ) and an encoder-side negative log-determinant penalty ( γ=0.2 ). The regularizer drives the layer imbalance Δ=W2⊤W2−W1W1⊤ towards Δ⋆=−(γ/λreg)INh=−2INh , corresponding to λbal=−2 . (B) Closed-form prediction from Appendix A.1.6 . (C) Unregularized control network initialized at λbal=−2 . This balance is conserved under gradient flow, providing a comparison with the balance selected dynamically in (A).
Figure 4: Effective rank and kernel distance of the NTK from initialization. Top row: Both metrics as functions of the scale τ and relative scale λ at step t=20,000 , without regularization ( wreg=0 ). Bottom row: The same metrics at the final checkpoint as functions of the scale τ and regularization weight wreg , with λ=0 fixed.
Training step
λ
0
100
1000
5000
10000
20000
−1.00
0.467
0.467
0.484
0.554
0.574
0.581
−0.75
0.467
0.467
0.487
0.563
0.587
0.595
−0.50
0.467
0.467
0.492
0.577
0.606
0.616
−0.25
0.467
0.469
0.500
0.596
0.627
0.639
Appendix
Table 3: Effective rank fraction across training steps for different values of layer imbalance λ .
Figure 5: Feature and kernel drift during self-supervised training. ImageNet-100, ViT-S/16 trained for 100 epochs with SACReg applied at the backbone. Left: feature drift from initialization, measured by 1−CKA at the CLS token after transformer blocks 3, 6, 9, and 12. Right: empirical NTK distance from initialization (solid) and from the previous checkpoint (dashed). Representations drift increasingly with depth, while the NTK moves substantially early in training and subsequently stabilizes, consistent with rich rather than lazy learning dynamics.
Figure 6: Schematic overview of λ -JEPA. Augmented views are mapped by the encoder ( fθ ) to hidden representations (h), then by ( pφ ) to embeddings (z). SACReg applies regularization directly in the hidden space to prevent per-view collapse, encouraging a more dispersed representation distribution (green) compared with training without SACReg (orange). Filled circles denote image centers; open circles denote individual views.
Method
Backbone
Ep.
DTD
Aircraft
Cars
CIFAR10
CIFAR100
Flowers
Food
Pets
Avg.
LeJEPA
ViT-S/16
100
69.4
43.8
43.4
89.3
71.1
84.0
73.6
75.3
68.7
DINO
ViT-S/16
100
69.3
55.2
57.0
93.8
78.6
89.3
76.7
85.9
75.7
iBOT
ViT-S/16
100
69.3
55.3
56.1
93.3
77.4
90.3
77.9
87.3
75.9
λ -JEPA
ViT-S/16
100
71.1
55.5
62.5
95.3
81.6
90.4
75.3
89.1
77.6
LeJEPA
ViT-S/16
300
70.2
47.6
50.3
92.1
74.4
85.8
75.7
79.1
71.9
DINO
ViT-S/16
300
71.5
59.8
66.2
95.0
81.1
92.0
79.6
89.6
79.4
Appendix
Table 4: Per-dataset linear-probe transfer accuracy (%) under the VISReg protocol; the last column is the average reported in Table 1 and 1 . ‡ Numbers reported by the VISReg paper; all other rows are evaluated by us under the matched protocol.
Figure 7: Effect of applying SACReg to the backbones across SSL objectives on ImageNet-100. We report representation statistics at the backbone h and projection space z , together with downstream linear probe and kNN accuracy. Open circles denote the baseline method and green circles the same method with SACReg added at the backbone. Each point is the mean over three seeds.
Figure 8: Effect of SACReg on view sensitivity and class-center separability. Left: mean class-level changes for six SSL methods on ImageNet-100, with arrows from the baseline representation to the same method with SACReg added at the backbone. Right: class-wise changes for λ -JEPA; 96 of 100 classes improve in class-center separability.
Readout at encoder output ( h )
LeJEPA
+ SIGReg at h
Positive-pair cosine
0.96
0.73
Negative-pair cosine
0.58
0.01
KL to N(0,1) (diagonal)
0.96
0.07
KL to N(0,I) (full covariance)
3.89
3.03
Epps–Pulley
652
504
RankMe /d
0.07
0.09
Appendix
Table 5: Adding SIGReg directly to the LeJEPA encoder output on ImageNet-100. The original projector losses are unchanged. All representation statistics are measured at the 512-dimensional encoder output.