After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning
Abstract
Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.
Figures & tables
| Arm | Critic | Value head | Value torso | Actor updates |
|---|---|---|---|---|
| X0 | learned | yes | yes | yes |
| X1 | learned | yes | no | yes |
| X5 | none | n/a | n/a | yes |
| Context | |||||
|---|---|---|---|---|---|
| H1 | 4.0 | 0.15 | 0.90 | 0.416 | 0.844 |
| H2 | 2.5 | 0.40 | 0.90 | 0.191 | 0.178 |
| H3 | 2.2 | 0.20 | 0.65 | 0.015 | 0.242 |
| Environment | Optimizer | Clipping | Role | Runs | Horizon | Primary summary |
|---|---|---|---|---|---|---|
| MinEx | Adam | global norm 0.5 | confirmatory | 900 | 196,608 env. steps | hazard / RMST |
| CleanUp-lite | SGD | none | confirmatory | 300 | 200 updates | paired RMST slope |
| MinEx | SGD | none | boundary pilot | 225 | 196,608 env. steps | raw events / actor activity |
| CleanUp-lite | Adam | global norm 0.5 | boundary | 300 | 200 updates | raw events / RMST |
| Arm | Scale | Events | Censored | Event rate | RMST | First-failure rate | Final cooperation |
|---|---|---|---|---|---|---|---|
| X0 | 0.25 | 2 | 58 | 0.033 | 0.9927 | 0.033 | 0.967 |
| X0 | 0.50 | 1 | 59 | 0.017 | 0.9986 | 0.083 | 0.883 |
| X0 | 1.0 | 30 | 30 | 0.500 | 0.8108 | 0.533 | 0.483 |
| X0 | 2.0 | 56 | 4 | 0.933 | 0.6003 | 0.967 | 0.017 |
| X0 | 4.0 | 60 | 0 | 1.000 | 0.4264 | 1.000 | 0.050 |
| X1 | 0.25 | 3 | 57 | 0.050 | 0.9958 | 0.050 | 0.933 |
| Context | Events/100 | RMST | Scale slope | SE |
|---|---|---|---|---|
| H1 | 60/100 | 0.665 | 1.935 | 0.191 |
| H2 | 44/100 | 0.767 | 2.776 | 0.311 |
| H3 | 45/100 | 0.865 | 1.957 | 0.258 |
| Arm | Scale | Events | Event rate | Mean RMST | Joint AUC | First-failure rate | Recovery rate |
|---|---|---|---|---|---|---|---|
| X0 | 0.25 | 0/20 | 0.00 | 200.0 | 13.764 | 0.00 | 0.00 |
| X0 | 0.50 | 0/20 | 0.00 | 200.0 | 13.751 | 0.00 | 0.00 |
| X0 | 1.0 | 0/20 | 0.00 | 200.0 | 13.651 | 0.00 | 0.00 |
| X0 | 2.0 | 0/20 | 0.00 | 200.0 | 13.672 | 0.00 | 0.00 |
| X0 | 4.0 | 11/20 | 0.55 | 143.9 | 8.389 | 0.90 | 0.40 |
| X1 | 0.25 | 0/20 | 0.00 | 200.0 | 13.770 | 0.00 | 0.00 |
| Checkpoint | X0 events | X1 events | X5 events |
|---|---|---|---|
| Base | 3/5 | 0/5 | 0/5 |
| Perturbation 1 | 4/5 | 0/5 | 0/5 |
| Perturbation 2 | 5/5 | 0/5 | 0/5 |
| Perturbation 3 | 3/5 | 0/5 | 0/5 |
| Perturbation 4 | 4/5 | 0/5 | 0/5 |
| Total | 19/25 | 0/25 | 0/25 |
| Mean RMST by scale | Mean joint AUC by scale | |||||||||
| Arm | 0.25 | 0.5 | 1 | 2 | 4 | 0.25 | 0.5 | 1 | 2 | 4 |
| X0 | 32.4 | 32.0 | 50.2 | 59.9 | 30.9 | 7.01 | 6.38 | 5.63 | 4.91 | 4.03 |
| X1 | 26.4 | 25.3 | 21.4 | 20.8 | 21.0 | 7.39 | 7.01 | 6.96 | 6.20 | 5.57 |
| X5 | 21.4 | 21.4 | 21.4 | 21.4 | 21.4 | 4.79 | 4.82 | 4.82 | 4.82 | 4.82 |
| Arm | Scale | Policy torso | Value torso | Value head | Global clip multiplier |
|---|---|---|---|---|---|
| X0 | 0.25 | 0.158 | 0.220 | 2.726 | 0.3625 |
| X0 | 0.50 | 0.155 | 0.743 | 8.705 | 0.1143 |
| X0 | 1.0 | 0.153 | 1.850 | 21.359 | 0.0466 |
| X0 | 2.0 | 0.152 | 4.069 | 46.913 | 0.0212 |
| X0 | 4.0 | 0.151 | 8.510 | 98.135 | 0.0102 |
| X1 | 0.25 | 0.158 | 0.000 | 2.726 | 0.3638 |
| Environment / direction | Margin | Estimate [95% CI] |
|---|---|---|
| CleanUp-lite arm-by-scale | Return / apples | [ ] |
| CleanUp-lite arm-by-scale | Pollution | [ ] |
| CleanUp-lite scale-4 X0–X1 | Return / apples | [ ] |
| CleanUp-lite scale-4 X0–X1 | Pollution | [ ] |
| MinEx scale-4 X0–X1 | Flow | [ ] |
| Component | MinEx –Adam | CleanUp-lite –SGD |
|---|---|---|
| Observation / actions | 14 / factorized | 82 / categorical 6 |
| Shared actor torso | Dense(128)–tanh–Dense(128)–tanh | Dense(128)–tanh–Dense(128)–tanh |
| Policy / value heads | three linear / scalar | six-way linear / scalar |
| Parallel environments | 16 (X0/X1); 8 (X5) | 8 |
| Rollout length | 128 (X0/X1); 512 (X5) | 64 |
| Epochs / minibatches | 4 / 4 | 4 / 4 |
| Contrast | Estimate | Cluster SE | Holm |
|---|---|---|---|
| X0 scale slope | 1.993 | 0.087 | |
| X0 slope X1 slope | 1.807 | 0.270 |