cs.LGMar 19, 2025

1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities

Authors: Kevin Wang, Ishaan Javali, Michał Bortkiewicz, Tomasz Trzciński, Benjamin Eysenbach

Organizations: Princeton University · Warsaw University of Technology · Warsaw University of Technology, Tooploox, IDEAS Research Institute

Abstract

Scaling up self-supervised learning has driven breakthroughs in language and vision, yet comparable progress has remained elusive in reinforcement learning (RL). In this paper, we study building blocks for self-supervised RL that unlock substantial improvements in scalability, with network depth serving as a critical factor. Whereas most RL papers in recent years have relied on shallow architectures (around 2 - 5 layers), we demonstrate that increasing the depth up to 1024 layers can significantly boost performance. Our experiments are conducted in an unsupervised goal-conditioned setting, where no demonstrations or rewards are provided, so an agent must explore (from scratch) and learn how to maximize the likelihood of reaching commanded goals. Evaluated on simulated locomotion and manipulation tasks, our approach increases performance on the self-supervised contrastive RL algorithm by 2×2\times - 50×50\times, outperforming other goal-conditioned baselines. Increasing the model depth not only increases success rates but also qualitatively changes the behaviors learned. The project webpage and code can be found here: https://wang-kevin3290.github.io/scaling-crl/.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Survival Reinforcement Learning: Toward Scalable Self-Supervised RL

    May 29, 2026Franki Nguimatsia-Tiofack, Fabian Schramm, Théotime Le Hellard +1Offline Reinforcement LearningScalable Robot Learning

  2. Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

    Jun 4, 2026Johan Obando-Ceron, Lu Li, Scott Fujimoto +3Multi-Turn Reinforcement LearningModel-Based Reinforcement Learning

  3. ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning

    Sep 28, 2026Nico Bohlinger, Jan PetersGoal-Conditioned Reinforcement LearningOffline Reinforcement Learning