cs.CVOct 7, 2026

Scalable Patch-Level Self-Supervised Learning

Authors: Maximilian Seitzer, Gabriele Trivigno, Antonín Vobecký, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Siméoni, Piotr Bojanowski

Organizations: Meta FAIR

Abstract

Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on 12×12\times less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning

    May 11, 2026Yang Shen, Yusen Cai, Weronika Hryniewska-Guzik +2Self-Supervised LearningSpatial Supervision

  2. Self-Supervised Visual Representation Learning: Pretrain-Finetuning or Joint Training?

    Jul 14, 2026Nusrat Munia, Tyler Ward, Nishat Nayla +2Self-Supervised LearningSelf-Supervised Representations

  3. Three Necessary Principles for Self-Supervised Visual Representation Learning

    Aug 8, 2026Nikos Giakoumoglou, Paschalis Giakoumoglou, Tania StathakiSelf-Supervised RepresentationsSelf-Supervised Learning