Modern LLM reinforcement learning (RL) workloads require a high-performance weight transfer system to scale training across heterogeneous compute resources. However, efficiently transferring terabyte-scale model weights across thousands of GPUs remains challenging because the system must accommodate clusters that dynamically scale up and down while keeping coordination, data movement, and storage overhead low. We introduce Reference-Oriented Storage (ROS), a new storage abstraction for RL weight transfer that exploits highly replicated model weights in place. ROS presents the illusion that certain versions of the model weights are stored and can be fetched on demand. Underneath, ROS does not physically store any copies of the weights; instead, it tracks the workers that hold these weights on GPUs for inference. Upon request, ROS directly uses them to serve reads. We build TensorHub, a production-quality system that instantiates the ROS idea with topology-aware transfer, model-parallel consistency, and fault tolerance. Evaluation shows that TensorHub saturates RDMA bandwidth and adapts to three distinct rollout workloads with minimal engineering effort. Specifically, TensorHub reduces total GPU stall time by up to 6.7x for standalone rollouts, accelerates weight updates for elastic rollouts by up to 4.8x, and cuts cross-datacenter rollout stall time by up to 19x. TensorHub has been deployed in ByteDance production to support cutting-edge RL training.
Figures & tables
Figure 1 . RL Workload with Diverse Rollout Patterns. Rollout workers stream in prompts and model weights and stream out responses. While prompts and responses are lightweight, the model weights can be several terabytes in size, necessitating an efficient transfer system.
Approach
Example
Topology Awareness
Elasticity
Minimal Coordination
Minimum Data Movement
No Extra Space
Collective Communication
NCCL ( Nvidia, [n. d.] )
✓
✓
✓
Point-to-Point Communication
UCX ( Shamis et al., 2015 )
✓
✓
✓
Distributed Storage
PS ( Li et al., 2014 ) , Ray object store ( Moritz et al., 2018 )
✓
✓
✓
TensorHub
✓
✓
✓
✓
✓
Table 1. Comparison of Different Weight Transfer Approaches.
Figure 2 . Reference-Oriented Storage Workflow. The server only operates on lightweight references. The bulk weight transfer is directly between clients.
Name
Description
open ( … ) -> ShardHandle
Acquire a shard handle. Optionally, declare the desired versions to retain for the retention protocol.
register (named_tensors)
Register weight tensors.
unregister ()
Unregister weight tensors.
publish (version)
Publish a reference to the registered tensors under the given version.
unpublish ()
Revoke the previously published reference.
replicate (version)
Replicate the given version of weights into registered tensors. Supports relative versions.
Table 2. TensorHub APIs. open() returns a handle, and the rest of the APIs are methods of this handle.
Figure 3. TensorHub Naming Scheme. Rollout-2 fetches version 3 by replicating a copy; after doing so, it also becomes a replica capable of serving subsequent replication requests.
Figure 4 . TensorHub Examples.
Figure 5 . Pipeline Replication. Worker-0 is the only source, while both Worker-1 and Worker-2 are requesting data. To scale throughput, TensorHub schedules a pipeline where Worker-2 reads partially replicated data on Worker-1.
Figure 6 . Sharding Consistency Example. Left: the physical request order lets replica-0’s shards observe different latest versions: shard:0 sees version 12 at T0 , while shard:1 sees version 13 after replica-1’s two shards publish it at T1 and T2 . Without coordination, this difference can cause SPMD control-flow divergence. Right: TensorHub serializes replica-0’s replicate requests as one transaction, so both shards observe version 12.
Figure 7 . Microbenchmark Results.
Model size
9B
36B
260B
mocked 1T
#shards
2
4
8
16
Shard size (GB)
10
19
34
66
Trainer #GPUs
16
16
64
768
Standalone #GPUs
8
8
16
256
Table 3. Training Workload Parameters.
Figure 8 . Weight Transfer Flows with Standalone. TensorHub does not require the Ray driver to orchestrate weight transfer. Each standalone rollout pulls weights on demand.
Figure 9 . Standalone Rollout Results. Ideally, only standalone rollouts need to stall due to their dependence on weights; trainers can proceed without waiting. TensorHub not only eliminates trainer stalls, but also keeps standalone stall time consistently lower.
Figure 10 . Weight Transfer Flows with Standalone and Elastic. Note that TensorHub only shows one possible data flow; an elastic rollout can also fetch weights from a trainer.
Figure 11 . Elastic Rollout Results.
Figure 12 . Cross-Datacenter Rollout Results. Note that the left figure shows only the rollout stall time; UCX incurs additional trainer stall time that is omitted here.