Nereus: Adaptive Parallelism for LLM Post-Training
Organizations: Aalto University · Shenzhen University of Advanced Technology · Zhejiang University
Abstract
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27 over OpenRLHF and by 1.10--1.47 over Verl across diverse clusters.
Figures & tables
| Approach | Representative systems | Timing ( when ) | Boundary ( what ) | Execution ( how ) | |
| A | Fixed-plan RL post-training execution | TRL ( von Werra et al., 2020 ) , DeepSpeed-Chat ( Yao et al., 2023 ) ; Verl ( Sheng et al., 2025 ) , AReaL ( Fu et al., 2025 ) , ReaL ( Mei et al., 2025 ) , PUZZLE ( Lei et al., 2024 ) , RLHFuse ( Zhong et al., 2025b ) , ROLL ( Wang et al., 2025 ) , OpenRLHF ( Hu et al., 2025 ) | none: plan fixed at startup | – | – |
| B | Checkpoint/restart | DCP ( Team, 2024b ) , MCP ( Team, 2024a ) , UCP ( Lian et al., 2025 ) , GCR ( Zeng et al., 2026 ) ; Gemini ( Wang et al., 2023 ) , ByteCheckpoint ( Wan et al., 2025 ) | on failure or restart | checkpoint state | checkpoint + restart |
| C | Single-model state management | Tenplex ( Wagenländer et al., 2024 ) , Oobleck ( Jang et al., 2023 ) | on a resource change | selective single-model state | host/GPU state transfer |
| D | Dynamic scheduling in a fixed pool | DynaRL ( Wang et al., 2026 ) | utilization and predicted throughput | per-component state | global allocation with per-component migration |
| Dependency-aligned elastic multi-model state management | Nereus | payback-based admission | dependency-aligned model-stage replica ( EMU ) | DAG of four primitives over GPU-direct links where available |
| Boundary | State managed | Planning unit | Cost |
|---|---|---|---|
| Checkpoint/restart | whole job | checkpoint | 836.74 s |
| Shard level | shards | shard | 66.43 s |
| Model-stage ( EMU ) | units | EMU | 6.52 s |
| Property | Cluster #1 | Cluster #2 | Cluster #3 |
|---|---|---|---|
| #Nodes / #GPUs | 128 / 1,024 | 64 / 256 | 8 / 64 |
| GPUs per node | 8 AMD MI250X GCDs (64 GB ) | 4 NVIDIA A100 64 GB | 8 NVIDIA H200 141 GB |
| Intra-node network | Infinity Fabric (400 GB/s) | NVLink 3.0 (600 GB/s) | NVLink 4.0 (900 GB/s) |
| CPUs per node | 1 64-core AMD EPYC 7A53 | 1 32-core Intel Xeon Platinum 8358 | 2 32-core Intel Xeon Platinum 8562Y+ |
| Host memory | 512 GB | 512 GB | 2 TB |
| Inter-node network | 4 HPE Cray Slingshot-11 (25+25 GB/s) | 4 HDR100 InfiniBand (12.5+12.5 GB/s) | 1 HDR InfiniBand (25+25 GB/s) |
| System | 4 8 | 8 16 | 16 32 | 32 64 |
|---|---|---|---|---|
| UCP † | 695.94 | 798.94 | 836.74 | 981.21 |
| Tenplex | 39.72 | 51.68 | 66.43 | 89.92 |
| Gemini † | 26.56 | 37.12 | 56.83 | 87.25 |
| Oobleck | 10.54 | 21.48 | 39.06 | 83.72 |
| Nereus | 2.45 | 5.58 | 6.52 | 8.48 |
| Type | Overlap | DynaRL-style | Nereus | Edges | ||
|---|---|---|---|---|---|---|
| GPUs | Succ. | Time | Succ. | Time | ||
| SM-CS | 32 | 62 | 8.48 | 100 | 9.80 | 1 |
| CM-SS | 48 | 56 | 8.93 | 100 | 10.80 | 2 |
| CM-CS | 48 | 34 | 10.44 | 100 | 12.60 | 2 |