Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.
Figures & tables
Figure 1: From discovery to maintenance. (A) Start from a verified cooperative actor. (B) X0 passes value gradients into the shared torso, X1 blocks them while retaining the critic, and X5 removes the critic. (C) Continued optimization varies optimizer and strategically equivalent reward scales. (D) Frozen evaluations track cooperation loss and recovery.
Arm
Critic
Value head
Value → torso
Actor updates
X0
learned
yes
yes
yes
X1
learned
yes
no
yes
X5
none
n/a
n/a
yes
Table 1: Matched routing arms. X0 and X1 share the critic, forward pass, and actor-update schedule, differing only in value-gradient access to the shared torso; X5 is a critic-free reference.
Context
H/L
c
q
mIC
mprofit
H1
4.0
0.15
0.90
0.416
0.844
H2
2.5
0.40
0.90
0.191
0.178
H3
2.2
0.20
0.65
0.015
0.242
Table 2: MinEx strategic contexts. All contexts support profitable, incentive-compatible cooperation; H3 has the narrowest incentive margin. mIC and mprofit denote the incentive-compatibility and profitability margins.
Environment
Optimizer
Clipping
Role
Runs
Horizon
Primary summary
MinEx
Adam
global norm 0.5
confirmatory
900
196,608 env. steps
hazard / RMST
CleanUp-lite
SGD
none
confirmatory
300
200 updates
paired RMST slope
MinEx
SGD
none
boundary pilot
225
196,608 env. steps
raw events / actor activity
CleanUp-lite
Adam
global norm 0.5
boundary
300
200 updates
raw events / RMST
Table 3: Experimental evidence matrix. Confirmatory cells test routing-specific scale effects across environments and optimizers; the opposite pair probes all-arm survival or failure.
Figure 2: MinEx / Adam confirmatory dynamics. (a) X0 RMST falls with scale across contexts. (b) X0 failures become earlier and more frequent. (c) At maximum scale, loss concentrates in X0; X1 approaches the censoring ceiling and X5 has no confirmed events. (d) Endpoint cooperation obscures failure timing.
Figure 3: CleanUp-lite / unclipped-SGD confirmatory dynamics. (a) The high-scale raster distinguishes transient failures, confirmed events, and censoring. (b,c) At maximum scale, survival deteriorates and confirmed failures appear only in X0. (d) Raw joint return reveals pre-event deterioration.
Figure 4: Mechanism diagnostics. (a) Direct value-to-torso gradients scale in X0 and vanish in X1. (b) Under global clipping, critic gradients still shrink X1 actor steps. (c) X5 fails under Adam but not SGD, revealing route-independent optimizer effects.
Figure 5: Local warm-start robustness in CleanUp-lite / SGD. (a) Across the base checkpoint and four local cooperative perturbations, X0 fails frequently only at high scale. (b) High-scale failures remain exclusive to X0; X1 and X5 remain event-free. Variants are perturbations, not independently trained policies.
Figure 6: Optimizer–environment boundary at reward scale 4. (a) Event rates show routing-specific failure, all-arm survival in MinEx –SGD, and all-arm failure in CleanUp-lite –Adam. (b) Wilson intervals confirm the pattern; direct value routing is neither necessary nor sufficient within this matrix.
Arm
Scale
Events
Censored
Event rate
RMST
First-failure rate
Final cooperation
X0
0.25
2
58
0.033
0.9927
0.033
0.967
X0
0.50
1
59
0.017
0.9986
0.083
0.883
X0
1.0
30
30
0.500
0.8108
0.533
0.483
X0
2.0
56
4
0.933
0.6003
0.967
0.017
X0
4.0
60
0
1.000
0.4264
1.000
0.050
X1
0.25
3
57
0.050
0.9958
0.050
0.933
Table 4: MinEx / Adam confirmatory survival summary. Higher reward scales increase X0 event rates and shorten its maintenance, whereas X1 remains near the censoring ceiling and X5 has no events. RMST is normalized by the 196,608-step horizon.
Context
Events/100
RMST
Scale slope
SE
H1
60/100
0.665
1.935
0.191
H2
44/100
0.767
2.776
0.311
H3
45/100
0.865
1.957
0.258
Table 5: X0 scale sensitivity appears in all three strategic contexts despite their different incentive margins. Each context pools five scales and 20 seeds per scale.
Arm
Scale
Events
Event rate
Mean RMST
Joint AUC
First-failure rate
Recovery rate
X0
0.25
0/20
0.00
200.0
13.764
0.00
0.00
X0
0.50
0/20
0.00
200.0
13.751
0.00
0.00
X0
1.0
0/20
0.00
200.0
13.651
0.00
0.00
X0
2.0
0/20
0.00
200.0
13.672
0.00
0.00
X0
4.0
11/20
0.55
143.9
8.389
0.90
0.40
X1
0.25
0/20
0.00
200.0
13.770
0.00
0.00
Table 6: CleanUp-lite / unclipped-SGD confirmatory summary. Confirmed failures occur only for high-scale X0; X1 and the critic-free X5 reference remain event-free across scales.
Checkpoint
X0 events
X1 events
X5 events
Base
3/5
0/5
0/5
Perturbation 1
4/5
0/5
0/5
Perturbation 2
5/5
0/5
0/5
Perturbation 3
3/5
0/5
0/5
Perturbation 4
4/5
0/5
0/5
Total
19/25
0/25
0/25
Table 7: Across the base checkpoint and four local variants, high-scale failures remain confined to X0; X1 and X5 have no events. Each cell uses five seeds.
Mean RMST by scale
Mean joint AUC by scale
Arm
0.25
0.5
1
2
4
0.25
0.5
1
2
4
X0
32.4
32.0
50.2
59.9
30.9
7.01
6.38
5.63
4.91
4.03
X1
26.4
25.3
21.4
20.8
21.0
7.39
7.01
6.96
6.20
5.57
X5
21.4
21.4
21.4
21.4
21.4
4.79
4.82
4.82
4.82
4.82
Table 8: CleanUp-lite / Adam boundary: all cells have 20/20 events. RMST and joint AUC show that identical event rates do not imply identical dynamics.
Arm
Scale
Policy → torso
Value → torso
Value → head
Global clip multiplier
X0
0.25
0.158
0.220
2.726
0.3625
X0
0.50
0.155
0.743
8.705
0.1143
X0
1.0
0.153
1.850
21.359
0.0466
X0
2.0
0.152
4.069
46.913
0.0212
X0
4.0
0.151
8.510
98.135
0.0102
X1
0.25
0.158
0.000
2.726
0.3638
Table 9: Production-path gradient audit. The direct value-to-torso gradient is nonzero and grows with scale in X0 but is exactly zero in X1. Norms describe the audited update and are not averaged treatment effects.
Environment / direction
Margin
Estimate [95% CI]
CleanUp-lite arm-by-scale
Return / apples
−0.0628 [ −0.1177,−0.0078 ]
CleanUp-lite arm-by-scale
Pollution
−0.00526 [ −0.00736,−0.00316 ]
CleanUp-lite scale-4 X0–X1
Return / apples
−0.0644 [ −0.1003,−0.0284 ]
CleanUp-lite scale-4 X0–X1
Pollution
−0.00554 [ −0.00678,−0.00431 ]
MinEx scale-4 X0–X1
Flow
−6.69×10−4 [ −1.26×10−3,−7.38×10−5 ]
Table 10: Centered local behavioral-margin estimates. Both CleanUp-lite directions point toward its cooperation boundary; in MinEx , only the flow margin under the high-scale X0–X1 direction is clearly adverse. Brackets give 95% batch-cluster intervals.
Component
MinEx –Adam
CleanUp-lite –SGD
Observation / actions
14 / factorized 5,3,3
82 / categorical 6
Shared actor torso
Dense(128)–tanh–Dense(128)–tanh
Dense(128)–tanh–Dense(128)–tanh
Policy / value heads
three linear / scalar
six-way linear / scalar
Parallel environments
16 (X0/X1); 8 (X5)
8
Rollout length
128 (X0/X1); 512 (X5)
64
Epochs / minibatches
4 / 4
4 / 4
Table 11: Matched implementation details for the confirmatory cells. The MinEx cell uses clipped Adam and the CleanUp-lite cell uses unclipped SGD; within each cell, X0 and X1 otherwise share the actor, critic computation, and update schedule.
Contrast
Estimate
Cluster SE
Holm p
X0 scale slope >0
1.993
0.087
<10−100
X0 slope > X1 slope
1.807
0.270
1.04×10−11
Table 12: Seed-clustered MinEx hazard contrasts. The X0 hazard increases with reward scale and more steeply than the matched X1 hazard.
Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute directly. Prior methods largely follow two approaches: independent factorized updates with centralized critics, which lack general joint-improvement guarantees without value decomposition assumptions, or alternating best-response updates, which can converge to suboptimal Nash Equilibria. In this paper, we show the joint policy gradient admits an exact decentralized decomposition of per-agent terms, each formed from per-agent score functions and decentralized critics. Based on this decomposition, we develop Agent-Chained Policy Optimization (ACPO), where actors are trained independently, with their updates together constituting a single step on the joint policy gradient. Central to this result is a serialized view of the simultaneous joint decision in which agents commit actions one at a time, each conditioning on a belief over preceding actions. The belief acts as the coordination mechanism which ties the independent per-agent updates into a joint gradient step. We evaluate ACPO on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo, where it outperforms strong baselines, with the gap widening as the number of agents grows.
Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara +4
Cooperative multi-agent reinforcement learning (MARL) benchmarks commonly emphasize aggregate outcomes such as return, success rate, or completion time. While essential, these metrics often fail to reveal how agents coordinate, particularly in settings where agents, tasks, and joint assignment choices scale combinatorially. We propose a coordination-aware evaluation perspective that supplements return with process-level diagnostics. We instantiate this perspective using STAT, a controlled commitment-constrained spatial task-allocation testbed that systematically varies agents, tasks, and environment size while holding observation access and task rules fixed. We evaluate six representative value-based MARL methods across varying levels of centralization. Our results show that similar return trends can reflect distinct coordination mechanisms, including differences in redundant assignment, assignment diversity, and task-completion efficiency. We find that in commitment-constrained task allocation, performance under scale is shaped not only by nominal action-space size, but also by assignment pressure, sparse decision opportunities, and redundant choices among interdependent agents. Our findings motivate coordination-aware evaluation as a necessary complement to return-based benchmarking for cooperative MARL.
Cooperative multi-agent reinforcement learning assumes each agent shares the same reward function and can be trained effectively using the Trust Region framework of single-agent. Instead of relying on other agents' actions, the independent actors setting considers each agent to act based only on its local information, thus having more flexible applications. However, in the sequential update framework, it is required to re-estimate the joint advantage function after each individual agent's policy step. Despite the practical success of importance sampling, the updated advantage function suffers from exponentially high variance problems, which likely result in unstable convergence. In this work, we first analyze the high variance advantage both empirically and theoretically. To overcome this limitation, we introduce a clipping objective to control the upper bounds of the advantage fluctuation in sequential updates. With the proposed objective, we provide a monotonic bound with sub-linear convergence to ε-Nash Equilibria. We further derive two new practical algorithms using our clipping objective. The experiment results on three popular multi-agent reinforcement learning benchmarks show that our proposed method outperforms the tested baselines in most environments. By carefully analyzing different training settings, our proposed method is highlighted with both stable convergence properties and the desired low advantage variance estimation. For reproducibility purposes, our source code is publicly available at https://github.com/giangbang/Low-Variance-Trust-Region-MARL.
Bang Giang Le, Viet Cuong Ta
Human Machine Interaction Laboratory, VNU University of Engineering and Technology, 144 Xuan Thuy, Cau Giay, 100000, Hanoi, Vietnam.