Scaling up self-supervised learning has driven breakthroughs in language and vision, yet comparable progress has remained elusive in reinforcement learning (RL). In this paper, we study building blocks for self-supervised RL that unlock substantial improvements in scalability, with network depth serving as a critical factor. Whereas most RL papers in recent years have relied on shallow architectures (around 2 - 5 layers), we demonstrate that increasing the depth up to 1024 layers can significantly boost performance. Our experiments are conducted in an unsupervised goal-conditioned setting, where no demonstrations or rewards are provided, so an agent must explore (from scratch) and learn how to maximize the likelihood of reaching commanded goals. Evaluated on simulated locomotion and manipulation tasks, our approach increases performance on the self-supervised contrastive RL algorithm by 2× - 50×, outperforming other goal-conditioned baselines. Increasing the model depth not only increases success rates but also qualitatively changes the behaviors learned. The project webpage and code can be found here: https://wang-kevin3290.github.io/scaling-crl/.
Figures & tables
Figure 1: Scaling network depth yields performance gains across a suite of locomotion, navigation, and manipulation tasks, ranging from doubling performance to 50× improvements on Humanoid-based tasks. Notably, rather than scaling smoothly, performance often jumps at specific critical depths (e.g., 8 layers on Ant Big Maze, 64 on Humanoid U-Maze), which correspond to the emergence of qualitatively distinct policies (see Section 4 ).
Figure 2: Architecture. Our approach integrates residual connections into both the actor and critic networks of the Contrastive RL algorithm. The depth of this residual architecture is defined as the total number of Dense layers across the residual blocks, which, with our residual block size of 4, equates to 4N .
Figure 3: Increasing depth results in new capabilities: Row 1 : A depth-4 agent collapses and throws itself toward the goal. Row 2 : A depth-16 agent walks upright. Row 3 : A depth-64 agent struggles and falls. Row 4 : A depth-256 agent vaults the wall acrobatically.
Figure 4: Scaling network width vs. depth . Here, we reflect findings from previous works ( Lee et al., 2024 ; Nauman et al., 2024b ) which suggest that increasing network width can enhance performance. However, in contrast to prior work, our method is able to scale depth, yielding more impactful performance gains. For instance, in the Humanoid environment, raising the width to 2048 (depth=4) fails to match the performance achieved by simply doubling the depth to 8 (width=256). The comparative advantage of scaling depth is more pronounced as the observational dimensionality increases.
Figure 5
Figure 7: Deeper networks unlock batch size scaling. We find that as depth increases from 4 to 64 in Humanoid, larger networks can effectively leverage batch size scaling to achieve further improvements.
Figure 8: We disentangle the effects of exploration and expressivity on depth scaling by training three networks in parallel: a “collector,” plus one deep and one shallow learner that train only from the collector’s shared replay buffer. In all three environments, when using a deep collector (i.e. good data coverage), the deep learner outperforms the shallow learner, indicating that expressivity is crucial when controlling for good exploration. With a shallow collector (poor exploration), even the deep learner cannot overcome the limitations of insufficient data coverage. As such, the benefits of depth scaling arise from a combination of improved exploration and increased expressivity working jointly.
Figure 9: Deeper Q-functions are qualitatively different. In the U4-Maze, the start and goal positions are indicated by the \mathbin{\vtop{\halign{#\cr\circledcirc\cr\bullet\crcr}}} and G symbols respectively, and the visualized Q values are computed via the L2 distance in the learned representation space, i.e., Q(s,a,g)=∥ϕ(s,a)−ψ(g)∥2 . The shallow depth 4 network (left) naively relies on Euclidean proximity, showing high Q values near the start despite a maze wall. In contrast, the depth 64 network (right) clusters high Q values at the goal, gradually tapering along the interior.
Figure 10: We visualize state-action embeddings from shallow (depth 4) and deep (depth 64) networks along a successful trajectory in the Humanoid task. Near the goal, embeddings from the deep network expand across a curved surface, while those from the shallow network form a tight cluster. This suggests that deeper networks may devote greater representational capacity to regions of the state space that are more frequently visited and play a more critical role in successful task completion.
Figure 11: Deeper networks exhibit improved generalization. (Top left) We modify the training setup of the Ant U-Maze environment such that start-goal pairs are separated by ≤3 units. This design guarantees that no evaluation pairs (Top right) were encountered during training, testing the ability for combinatorial generalization via stitching. (Bottom) Generalization ability improves as network depth grows from 4 to 16 to 64 layers.
Figure 12: Testing the limits of scale. We extend the results from Figure 1 by scaling networks even further on the challenging Humanoid maze environments. We observe continued performance improvements with network depths of 256 and 1024 layers on Humanoid U-Maze. Note that for the 1024-layer networks, we observed the actor loss exploding at the onset of training, so we maintained the actor depth at 512 while using 1024-layer networks only for the two critic encoders.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Scaled CRL (Ours) outperforms baselines CRL (original), SAC, SAC+HER, TD3+HER, GCSL, and GCBC in 8 out 10 environments.
Figure 13: Depth scaling yields limited gains for SAC, SAC+HER, TD3+HER, GCSL, and GCBC.
Task
Dim
D = 4
D = 64
Imprv.
Arm Binpick Hard
17
38±4
219±15
5.7×
Arm Push Easy
308±33
762±30
2.5×
Arm Push Hard
171±11
410±13
2.4×
Ant U4-Maze
29
11.4±4.1
286±36
25×
Ant U5-Maze
0.97±0.7
61±18
63×
Ant Big Maze
61±20
441±25
7.3×
Appendix
Table 1: Increasing network depth (depth D=4→64 ) increases performance on CRL ( Figure 1 ). Scaling depth exhibits the greatest benefits on tasks with the largest observation dimension ( Dim ).
Figure 14: Our approach successfully scales depth in offline GCBC on antmaze-medium-stitch (OGBench). In contrast, scaling depth for BC ( antmaze-giant-navigate , expert SAC data) and for both online ( FetchPush ) and offline QRL ( pointmaze-giant-stitch , OGBench) yield negative results.
Figure 15: Performance of depth scaling on CRL augmented with quasimetric architectures (CMD-1).
Figure 16: (Left) Layer Norm is essential for scaling depth. (Right) Scaling with ReLU activations leads to worse performance compared to Swish activations.
Steps to reach ≥ 200 success Depth 4 16 32 With – 50 42 Without – 64 54
Steps to reach ≥ 400 success Depth 4 16 32 With – 62 48 Without – 75 64
Steps to reach ≥ 600 success Depth 4 16 32 With – 77 67 Without – – 77
Appendix
Table 2: Integrating hyperspherical normalization in our architecture enhances the sample efficiency of depth scaling.
Figure 17: L2 norms of residual activations in networks with depths of 32, 64, 128, and 256.
Figure 18: To evaluate the scalability of our method in the offline setting, we scaled model depth on OGBench ( Park et al., 2024 ) . In two out of three environments, performance drastically declined as depth scaled from 4 to 64, while a slight improvement was seen on antmaze-medium-stitch-v0. Successfully adapting our method to scale offline GCRL is an important direction for future work.
Figure 19: The scaling results of this paper are demonstrated on the JaxGCRL benchmark, showing that they replicate across a diverse range of locomotion, navigation, and manipulation tasks. These tasks are set in the online goal-conditioned setting where there are no auxiliary rewards or demonstrations. Figure taken from ( Bortkiewicz et al., 2024 ) .
Figure 20: Scaling behavior for humanoid in two different python environments: MJX=3.2.3, Brax=0.10.5 and MJX=3.2.6, Brax=0.10.1 (ours) version of JaxGCRL. Scaling depth improves the performance significantly for both versions. In the environment we used, training requires fewer environment steps to reach a marginally better performance than in other Python environment.
Environment
Depth 4
Depth 8
Depth 16
Depth 32
Depth 64
Humanoid
1.48 ± 0.00
2.13 ± 0.01
3.40 ± 0.01
5.92 ± 0.01
10.99 ± 0.01
Ant Big Maze
2.12 ± 0.00
2.77 ± 0.00
4.04 ± 0.01
6.57 ± 0.02
11.66 ± 0.03
Ant U4-Maze
1.98 ± 0.27
2.54 ± 0.01
3.81 ± 0.01
6.35 ± 0.01
11.43 ± 0.03
Ant U5-Maze
9.46 ± 1.75
10.99 ± 0.02
16.09 ± 0.01
31.49 ± 0.34
46.40 ± 0.12
Ant Hardest Maze
5.11 ± 0.00
6.39 ± 0.00
8.94 ± 0.01
13.97 ± 0.01
23.96 ± 0.06
Arm Push Easy
9.97 ± 1.03
11.02 ± 1.29
12.20 ± 1.43
14.94 ± 1.96
19.52 ± 1.97
Appendix
Table 3: Wall-clock time (in hours) for Depth 4, 8, 16, 32, and 64 across all 10 environments.
Depth
Time (h)
4
3.23 ± 0.001
8
4.19 ± 0.003
16
6.07 ± 0.003
32
9.83 ± 0.006
64
17.33 ± 0.003
128
32.67 ± 0.124
Appendix
Table 4: Total wall-clock time (in hours) for training from Depth 4 up to Depth 1024 in the Humanoid U-Maze environment.
Environment
Scaled CRL
SAC
SAC+HER
TD3
GCSL
GCBC
Humanoid
11.0 ± 0.0
0.5 ± 0.0
0.6 ± 0.0
0.8 ± 0.0
0.4 ± 0.0
0.6 ± 0.0
Ant Big Maze
11.7 ± 0.0
1.6 ± 0.0
1.6 ± 0.0
1.7 ± 0.0
1.5 ± 0.3
1.4 ± 0.1
Ant U4-Maze
11.4 ± 0.0
1.2 ± 0.0
1.3 ± 0.0
1.3 ± 0.0
0.7 ± 0.0
1.1 ± 0.1
Ant U5-Maze
46.4 ± 0.1
5.7 ± 0.0
6.1 ± 0.0
6.2 ± 0.0
2.8 ± 0.1
5.6 ± 0.5
Ant Hardest Maze
24.0 ± 0.0
4.3 ± 0.0
4.5 ± 0.0
5.0 ± 0.0
2.1 ± 0.6
4.4 ± 0.5
Arm Push Easy
19.5 ± 0.6
8.3 ± 0.0
8.5 ± 0.0
8.4 ± 0.0
6.4 ± 0.1
8.3 ± 0.3
Appendix
Table 5: Wall-clock training time comparison of our method vs. baselines across all 10 environments.
Environment
SAC
Scaled CRL (Depth 64)
Humanoid
0.46
6.37
Ant Big Maze
1.55
0.00
Ant U4-Maze
1.16
0.00
Ant U5-Maze
5.73
0.00
Ant Hardest Maze
4.33
0.45
Arm Push Easy
8.32
1.91
Appendix
Table 6: Wall-clock time (in hours) for our approach to surpass SAC’s final performance. As shown, our approach surpasses SAC performance in less wall-clock time in 7 out of 10 environments. The N/A* entries are because in those environments, scaled CRL doesn’t outperform SAC.