TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
Organizations: ByteDance Seed · University of Wisconsin–Madison
Abstract
Modern LLM reinforcement learning (RL) workloads require a high-performance weight transfer system to scale training across heterogeneous compute resources. However, efficiently transferring terabyte-scale model weights across thousands of GPUs remains challenging because the system must accommodate clusters that dynamically scale up and down while keeping coordination, data movement, and storage overhead low. We introduce Reference-Oriented Storage (ROS), a new storage abstraction for RL weight transfer that exploits highly replicated model weights in place. ROS presents the illusion that certain versions of the model weights are stored and can be fetched on demand. Underneath, ROS does not physically store any copies of the weights; instead, it tracks the workers that hold these weights on GPUs for inference. Upon request, ROS directly uses them to serve reads. We build TensorHub, a production-quality system that instantiates the ROS idea with topology-aware transfer, model-parallel consistency, and fault tolerance. Evaluation shows that TensorHub saturates RDMA bandwidth and adapts to three distinct rollout workloads with minimal engineering effort. Specifically, TensorHub reduces total GPU stall time by up to 6.7x for standalone rollouts, accelerates weight updates for elastic rollouts by up to 4.8x, and cuts cross-datacenter rollout stall time by up to 19x. TensorHub has been deployed in ByteDance production to support cutting-edge RL training.
Figures & tables
| Approach | Example | Topology Awareness | Elasticity | Minimal Coordination | Minimum Data Movement | No Extra Space |
| Collective Communication | NCCL ( Nvidia, [n. d.] ) | ✓ | ✓ | ✓ | ||
| Point-to-Point Communication | UCX ( Shamis et al., 2015 ) | ✓ | ✓ | ✓ | ||
| Distributed Storage | PS ( Li et al., 2014 ) , Ray object store ( Moritz et al., 2018 ) | ✓ | ✓ | ✓ | ||
| TensorHub | ✓ | ✓ | ✓ | ✓ | ✓ | |
| Name | Description |
| open ( ) -> ShardHandle | Acquire a shard handle. Optionally, declare the desired versions to retain for the retention protocol. |
| register (named_tensors) | Register weight tensors. |
| unregister () | Unregister weight tensors. |
| publish (version) | Publish a reference to the registered tensors under the given version. |
| unpublish () | Revoke the previously published reference. |
| replicate (version) | Replicate the given version of weights into registered tensors. Supports relative versions. |
| Model size | 9B | 36B | 260B | mocked 1T |
| #shards | 2 | 4 | 8 | 16 |
| Shard size (GB) | 10 | 19 | 34 | 66 |
| Trainer #GPUs | 16 | 16 | 64 | 768 |
| Standalone #GPUs | 8 | 8 | 16 | 256 |