Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on 12× less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.
Figures & tables
Figure 1: We introduce JEM , a student-teacher SSL objective grounded in information theory. Following the multi-view principle (a), it maximizes the information a student representation retains about what views share. Average scores over semantic, panoptic and instance segmentation at different scales (b) show that JEM scales effectively and consistently outperforms DINOv2. (c) JEM surpasses methods using patch- and/or image-level objectives on dense tasks while remaining competitive on classification. JEM learns strong representations, is principled and stable.
Figure 2: From multi-view InfoMax to JEM . Maximizing the variational lower bound on the multi-view InfoMax objective in Eq. 2 yields two requirements: cross-view alignment between teacher and student distributions p and q , and informative, non-collapsed targets T∼q(⋅∣XT) . JEM implements this with three loss terms: the KL term Lalign aligns student patch ℓ with matching teacher target m(ℓ) ; Rglobal maintains per-patch mutual information through marginal and conditional entropy; and Rstruct seeks to control the total correlation of the targets \TC(T) by preserving the geometry of teacher features of layer k in the final student representation. Please zoom for details.
Figure 3: View construction.
Table 1: Ablation study. We isolate the contributions of view construction (a), the composition of the patch-level KL alignment loss (b), and the terms in the complete objective (c). We train ViT-L models on ImageNet-22k for 250k steps and report attention-probe accuracy on IN1k and segmentation mIoU on ADE20k. Tables compare the full model with variants removing (w/o) individual components.
Figure 5
Class.
Sem. Seg.
Inst. Seg.
Pan. Seg.
Depth
Model
Size
Dataset
Tokens (T)
IN1k
ADE
VOC
City.
COCO
ADE
COCONut
NYU
KITTI
MAE
L
IN-1k
0.40
78.4
31.6
66.1
53.1
6.0
6.5
9.7
0.489
2.903
MSN
L
IN-1k
0.43
76.3
24.3
55.0
46.4
3.2
4.4
6.7
0.656
3.696
data2vec 2.0
L
IN-1k
0.60
79.5
27.4
52.7
48.6
2.1
4.8
7.1
0.512
2.901
CAPI
L
IN-1k
2.10
82.9
31.9
66.0
50.6
10.7
12.7
14.7
0.361
2.604
DINOv2*
L
IN-1k
0.23
83.8
45.7
83.3
66.5
12.1
18.0
22.1
0.408
2.718
Table 2: Comparison to prior SSL methods . Classification reports IN-1k attention-probe accuracy. Semantic, instance, and panoptic segmentation results are mIoU, mAP, and PQ, respectively. Depth reports RMSE (lower better). Tokens is the approximate number of tokens seen during training in trillions. ‘*’ denotes the DINOv2 algorithm trained with the JEM setup for 125k steps on IN-1k and 500k steps on IN-22k or A140M. † denotes backbone training or fine-tuning above 256 px.
Figure 6: Scaling model size. Performance as a function of approximate training compute, averaged over datasets for each benchmark type. DINOv2* and JEM are trained by us for 500k steps on A140M. The dashed purple lines show DINOv3 ViT-7B.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Scaling across individual benchmarks. Performance as a function of approximate training compute. DINOv2* and JEM are trained by us on A140M for 500k steps. The purple dashed lines show DINOv3 ViT-7B. ADE20k is evaluated with both semantic segmentation (mIoU) and panoptic segmentation (PQ) probes.
Figure 8: Student-teacher graphical models. (a) shows the generic model with the form of T not specified; (b) shows the element-wise model with additional factorization assumptions. The views (XS,XT) are sampled jointly given input X . Solid arrows describe the sampling model, with R=f(XS) and T∼p(⋅∣XT) ; double circles mark deterministic nodes. Dashed arrows denote the variational predictors q(T∣R) and q(T(ℓ)∣R(ℓ)) . In (b), ℓ indexes matched patch/layer pairs, with teacher targets relabeled accordingly and R(ℓ)=f(ℓ)(XS) . The plate represents the factorizations p(T∣XT)=∏ℓ=1Lp(T(ℓ)∣XT) and q(T∣R)=∏ℓ=1Lq(T(ℓ)∣R(ℓ)) .
Model
Citation
Arch. / Variant
Training Data
Params.
External Checkpoint
MAE
He et al. (2022)
ViT-L/16
IN-1k
303M
facebook/vit-mae-large
MSN
Assran et al. (2022)
ViT-L/16
IN-1k
303M
vitl16_600ep
data2vec 2.0
Baevski et al. (2023)
ViT-L/16
IN-1k
303M
large_imagenet
I-JEPA
Assran et al. (2023)
ViT-g/16
IN-22k
1,011M
IN22K-vit.g.16-600e
DINOv2 (reg.)
Oquab et al. (2024)
ViT-g/14
LVD-142M
1,136M
dinov2_vitg14_reg
CAPI
Darcet et al. (2025)
ViT-L/14
IN-1k
303M
capi_vitl14_in1k
Appendix
Table 3: External checkpoints evaluated in , grouped by method and ordered by first public appearance.
Model
Arch.
Stage / input
Batch size
Updates
Student views
Tokens / sample
Tokens (T)
MAE
L/16
Pretraining
4,096
≈500,456
1×2242
196
0.402
MSN
L/16
Pretraining
1,024
≈750,684
1×2242+10×962
556
0.427
data2vec 2.0
L/16
Pretraining
256
750,000
16 masks×2242
3,136
0.602
CAPI
L/14
Pretraining
16,384
500,000
1×2242
256
2.097
DINOv2*
L/16
Pretraining
2,048
125,000
2×2562+8×1122
904
0.231
JEM
L/16
Pretraining
2,048
125,000
2×2562+8×1122
904
0.231
Appendix
Table 4: Estimated training tokens for the checkpoints in . V-JEPA video inputs use temporal tubelets of two frames. Student views list crop counts and spatial dimensions in pixels, with frame counts for video. Batch size counts images or clips before crop and mask replication. Multi-stage models include a total row; total token counts are in trillions ( 1012 ). Details and assumptions are given in .
Table 5: Shared architectures and training hyperparameters for JEM and DINOv2 baselines . Entries spanning the three model columns are shared across all model sizes.
Hyperparameter
Model size
ViT-L
ViT-g
ViT-7B
Default training parameters
Teacher-crop scale
[0.32, 1.0]
Number of categories
4096
Student and teacher temperature
0.1
Global mask probability / patch ratio
0.8 / 0.65
Appendix
Table 6: JEM training hyperparameters . Entries spanning the three model columns are shared across all model sizes. Shared architectures and training hyperparameters are given in . Individual ablation settings are listed in .
Table 7: Ablation settings. Experiment-specific settings for the ablations in . All experiments use the shared ablation training settings in .
Hyperparameter
Model size
ViT-L
ViT-g
ViT-7B
Global-crop scale
[0.32,1.00]
Local-crop scale
[0.05,0.32]
Global mask probability
0.5
Mask-ratio range
[0.1,0.5]
DINO prototypes
65,536
131,072
262,144
Appendix
Table 8: DINOv2 baseline training hyperparameters . Entries spanning all model columns are shared across model sizes. Shared architectures and training hyperparameters are given in . Loss weights are configured coefficients.
Figure 9: PCA visualization of final-layer patch representations of a JEM ViT-g model. For each of the 20 images, we show the input and its RGB projection onto the first three principal components computed individually on each image.
Figure 10: Cosine similarity between final-layer patch representations and the center-patch representation of a JEM ViT-g model. Red markers identify the reference patch. The color scale interpolates between similarity values of 0 (full blue) and 1 (full yellow).
Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationships among object parts. To address this limitation, we introduce Spatial Prediction (SP), a spatially aware pretext regression task that predicts the relative position and scale between a pair of disentangled local views from the same image. By modeling part-to-part relationships in a continuous geometric space, SP encourages representations to capture fine-grained spatial dependencies beyond invariant categorical semantics, thereby learning the compositional structure of visual scenes. SP is implemented as a decoupled plug-in and can be seamlessly integrated into diverse SSL frameworks. Extensive experiments show consistent improvements across image recognition, fine-grained classification, semantic segmentation, and depth estimation, as well as substantial gains in out-of-distribution robustness for object recognition. To evaluate spatial reasoning, we introduce (1) a position and scale prediction task on image patch pairs and (2) a jigsaw understanding task requiring patch reordering and recognition after reconstruction. Strong performance on these tasks indicates improved spatial structure and geometric awareness. Overall, explicitly modeling spatial information provides an effective inductive bias for SSL, leading to more structured representations and better generalization. Code and models will be released.
Yang Shen, Yusen Cai, Weronika Hryniewska-Guzik +2
Nanyang Technological University, Singapore · Warsaw University of Technology, Poland
Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approach for self-supervised learning (SSL): a pretraining stage on unlabeled data followed by a finetuning stage on labeled data. While this pipeline has demonstrated extreme effectiveness, the interaction between self-supervised and supervised learning objectives remains insufficiently understood. In this work, we systematically investigate whether jointly optimizing the self-supervised and supervised objectives during training provides a better alternative. We compare two training paradigms: (1) the aforementioned pretraining followed by finetuning (PFT) and (2) joint training (JT), where self-supervised and supervised losses are optimized simultaneously in the same network. Across eight representative SSL methods and diverse computer vision tasks on natural, medical, crisis response, and remote sensing data, we evaluate performance under varying percentages of labeled data. Our results reveal that the relative effectiveness of PFT and JT depends strongly on the task at hand, the availability of labeled data, and the complexity of the domain. We find that JT consistently improves data and training efficiency while being robust in low-label settings, while PFT is more reliable in more specialized domains. We further analyze representation quality, robustness, and cross-domain generalization, providing new insights into how self-supervised and supervised objectives interact during optimization. We establish a comprehensive empirical benchmark for hybrid SSL-based semi-supervised learning and offer practical guidance for selecting appropriate training strategies across diverse vision applications.
Nusrat Munia, Tyler Ward, Nishat Nayla +2
Department of Computer Science University of Kentucky Lexington, KY 40506, USA · Kentucky Geological Survey University of Kentucky Lexington, KY 40506, USA
We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. We formalize these as the observation, prediction, and regularization principles and prove (i) that combining observation and prediction without regularization admits the constant encoder as a global minimizer under negative-free alignment; (ii) that the two objectives are gradient-complementary and structurally non-conflicting at the encoder output; and (iii) that the momentum encoder converges to the same fixed point as the online encoder and provides no collapse guarantee at convergence. Contrastive alignment provides only self-limiting collapse resistance, formalized via an explicit gradient-decay argument. Dropping prediction withholds the spatial training signal by construction; dropping observation forfeits cross-view semantic invariance by construction; at the scale we study, no pair substitutes for the third. Every major self-supervised method is a special case of a single unified energy decomposition. We pair every theoretical claim with a controlled experiment, including a patch-retrieval evaluation for the spatial consequence of prediction.