Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27× over OpenRLHF and by 1.10--1.47× over Verl across diverse clusters.
Figures & tables
Figure 1 . Nereus overview. (a) 1 The controller replans on drift and admits the transition to S∗ only if the current plan is infeasible or the savings repay the transition cost (§ 4 ). 2 Both plans are represented as EMU s, one per model-stage replica, with TP/PP inside the unit and DP as the replica count (§ 5 ). 3 The transition engine compiles the plan difference into a global transition DAG, adding resource-dependency edges when transient GPU overlap blocks an acquisition (§ 6 ). (b) Sequence-length growth reshards the five model-stages, with Merge in generation and inference, and Split in training ( S1→S2 ). A change in the GPU count removes an actor-training replica with Destroy and adds critic replicas with Extend ( S2→S3 ).
Figure 2 . Multi-model, multi-stage workflow of RL post-training with Proximal Policy Optimization (PPO).
Figure 3 . Frequent GPU availability fluctuations in shared cloud environments ( p3.2xlarge nodes) over ten hours.
Figure 4 . Generated sequence length grows during training (a), changing which measured training plan is fastest (b). Missing data points indicate out-of-memory (OOM) errors.
Approach
Representative systems
Timing ( when )
Boundary ( what )
Execution ( how )
A
Fixed-plan RL post-training execution
TRL ( von Werra et al., 2020 ) , DeepSpeed-Chat ( Yao et al., 2023 ) ; Verl ( Sheng et al., 2025 ) , AReaL ( Fu et al., 2025 ) , ReaL ( Mei et al., 2025 ) , PUZZLE ( Lei et al., 2024 ) , RLHFuse ( Zhong et al., 2025b ) , ROLL ( Wang et al., 2025 ) , OpenRLHF ( Hu et al., 2025 )
none: plan fixed at startup
–
–
B
Checkpoint/restart
DCP ( Team, 2024b ) , MCP ( Team, 2024a ) , UCP ( Lian et al., 2025 ) , GCR ( Zeng et al., 2026 ) ; Gemini ( Wang et al., 2023 ) , ByteCheckpoint ( Wan et al., 2025 )
on failure or restart
checkpoint state
checkpoint + restart
C
Single-model state management
Tenplex ( Wagenländer et al., 2024 ) , Oobleck ( Jang et al., 2023 )
on a resource change
selective single-model state
host/GPU state transfer
D
Dynamic scheduling in a fixed pool
DynaRL ( Wang et al., 2026 )
utilization and predicted throughput
per-component state
global allocation with per-component migration
Dependency-aligned elastic multi-model state management
Nereus
payback-based admission
dependency-aligned model-stage replica ( EMU )
DAG of four primitives over GPU-direct links where available
Table 1 . Approaches to online RL post-training adaptation along the three design decisions of § 3.3 : when to adapt (timing), what state to reuse (the boundary), and how to transition. Nereus uses a dependency-aligned model-stage boundary.
Figure 5 . Core EMU primitives for online adaptation: Split (a) and Merge (b) reshard state by converting between tightly coupled TP/PP structure and loosely coupled replicas, while Extend (c) and Destroy (d) scale resources by creating or removing replicas. For the example in § 5.1 , Extend adds replicas for critic training, and Extend followed by Merge increases TP for actor generation without changing its DP degree.
Boundary
State managed
Planning unit
Cost
Checkpoint/restart
whole job
checkpoint
836.74 s
Shard level
shards
shard
66.43 s
Model-stage ( EMU )
units
EMU
6.52 s
Table 2 . State boundaries for online adaptation, with the measured resource-scaling cost for an 8B actor/critic workload from 16 to 32 GPUs on Cluster #3 (the 16 → 32 column of Tab. 4 ). The checkpoint row is measured with UCP ( Lian et al., 2025 ) and the shard-level row with Tenplex ( Wagenländer et al., 2024 ) .
Figure 6 . Constructing the global transition DAG from S to S∗ : Nereus first translates each model-stage’s change into a local DAG, then adds cross-model-stage dependency edges under transient GPU overlap.
Property
Cluster #1
Cluster #2
Cluster #3
#Nodes / #GPUs
128 / 1,024
64 / 256
8 / 64
GPUs per node
8 × AMD MI250X GCDs (64 GB )
4 × NVIDIA A100 64 GB
8 × NVIDIA H200 141 GB
Intra-node network
Infinity Fabric (400 GB/s)
NVLink 3.0 (600 GB/s)
NVLink 4.0 (900 GB/s)
CPUs per node
1 × 64-core AMD EPYC 7A53
1 × 32-core Intel Xeon Platinum 8358
2 × 32-core Intel Xeon Platinum 8562Y+
Host memory
512 GB
512 GB
2 TB
Inter-node network
4 × HPE Cray Slingshot-11 (25+25 GB/s)
4 × HDR100 InfiniBand (12.5+12.5 GB/s)
1 × HDR InfiniBand (25+25 GB/s)
Table 3 . Testbed GPU clusters. GCD denotes a graphics compute die, counted as one GPU. Bandwidths are nominal, and inter-node bandwidth is listed per link as send+receive.
Figure 7 . Llama-3.1-8B PPO performance at fixed GPU budgets on Cluster #2.
Figure 8 . Strong scalability on Cluster #1.
Figure 9 . PPO throughput across 8B–70B models and GPU budgets on Cluster #3.
Figure 10 . Cost-model fit to measured training-stage latency (a) and target-plan selection time (b). Annotations in (a) list the (DP, TP, PP) tuples for two example global plans.
Figure 11 . Controller policies under held-out sequence-length and GPU-supply drift: step latency, gap to the best plan at each step, and sensitivity to the confidence margin.
System
4 → 8
8 → 16
16 → 32
32 → 64
UCP †
695.94
798.94
836.74
981.21
Tenplex
39.72
51.68
66.43
89.92
Gemini †
26.56
37.12
56.83
87.25
Oobleck
10.54
21.48
39.06
83.72
Nereus
2.45
5.58
6.52
8.48
Table 4. Cost (s) of doubling the GPUs for the 8B actor/critic workload on Cluster #3. † Also persists state.
Figure 12 . Parallelism-resharding cost on Cluster #3. MCP also persists state. Missing bars indicate OOM in all systems.
Figure 13 . Transition planning across 32–1,024 GPUs.
Type
Overlap
DynaRL-style
Nereus
Edges
GPUs
Succ.
Time
Succ.
Time
SM-CS
32
62
8.48
100
9.80
1
CM-SS
48
56
8.93
100
10.80
2
CM-CS
48
34
10.44
100
12.60
2
Table 5 . Coordination under three transition types: success rate (%), successful-trial time (s), and added resource-dependency edges. DynaRL-style uses per-component migration with independent local DAGs. SM-CS, CM-SS, and CM-CS denote same-model cross-stage, cross-model same-stage, and cross-model cross-stage overlap.
Figure 14 . Step latency on Cluster #3 for ReMax and GRPO (eight samples per prompt) and asynchronous RL.
Figure 15 . Training behavior over the first 50 steps: reward against wall-clock time (left), reward and PPO KL estimate against step (middle, right).
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves 2.17--2.79× the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9% for Qwen3-8B and 36.3% for Qwen3-30B-A3B.
Modern large language model (LLM) training is inherently dynamic: resource fluctuations, RLHF phase shifts, and cluster elasticity continually reshape the optimal parallelism layout, posing a significant challenge to existing training frameworks built around a static execution model. We present DynaTrain, a distributed training system for sub-second, online reconfiguration across arbitrary multi-dimensional parallelism. At its core, we propose a Virtual Parameter Space (VPS) abstraction that unifies all distributed training states under one logical coordinate space, turning any parallelism configuration into a deterministic mapping and collapsing complex transition into manageable geometric intersections. On top of VPS, a state routing-and-transition layer executes rank-local transfers under a memory-aware, deadlock-free schedule, and an Elastic Device Manager overlaps new-world construction with ongoing training to mask topology-change cost. On dense and MoE models up to 235B parameters, DynaTrain reconfigures a 70B dense model in under 2s and a 235B MoE model in 4.36s, outperforming state-of-the-art checkpoint-based and elastic systems by up to three orders of magnitude while preserving correctness.
Yuanqing Wang, Yuchen Zhang, Hao Lin +9
Peking University · Infinigence AI · Institute of Computing Technology, CAS +2
We present JigsawRL, a cost-efficient framework that explores Pipeline Multiplexing as a new dimension of RL parallelism. JigsawRL decomposes each pipeline into a Sub-Stage Graph that exposes the intra-stage and inter-worker imbalance hidden by stage-level systems. On this abstraction, JigsawRL resolves multiplexing interference through dynamic resource allocation, eliminates fragmented utilization by migrating long-tail rollouts across workers, and formulates their coordination as a graph scheduling problem solved with a look-ahead heuristic. On 4-64 H100/A100 GPUs across different agentic RL pipelines and models, JigsawRL achieves up to 1.85x throughput over Verl on synchronous RL, 1.54x over StreamRL and AReaL on asynchronous RL, and supports heterogeneous pipelines with moderate latency trade-off.