Constrained multi-agent control requires more than predicting rewarding actions: an action can cease to be executable as contact windows, shared capacity, and deadlines change. We introduce VERA, a centralized-training, decentralized-execution framework that separates feasibility estimation from credit assignment. Each actor predicts a five-dimensional verifiable feasibility representation (VFR). After an action is proposed, exact action-conditioned margins available only during training supervise that representation, while a counterfactual group-relative advantage (CGRA) ranks candidate representation-action pairs. Execution uses one actor pass and no privileged state. In a dynamic space-air-ground integrated network (SAGIN), VERA obtains 55.33% +/- 3.60% success with 0.45% +/- 0.81% coverage violation, within 1.33 percentage points of a privileged-mask reference. With rewards matched over ten paired seeds, VERA improves success over the strongest baseline by 8.74 percentage points (p=0.023) and reduces violation by 52.19 percentage points (p=5.7e-8). A ten-seed 4-by-2 factorial attributes a 14.16-16.48 percentage-point gain to CGRA across handcrafted, learned, random, and latent representations; evaluation on seven unseen topologies preserves a 24.33-30.02 percentage-point advantage over multi-agent proximal policy optimization. From 10 to 40 users, success remains 50.1-53.8%, and VFR adds only 0.026 ms to a central processing unit (CPU) actor step. Cross-domain tests further identify the governing condition: counterfactual credit succeeds when candidate scores respect shared constraints and fails under incompatible reward geometries. These results establish action-conditioned feasibility as an auditable training interface and counterfactual credit as a geometry-dependent optimization mechanism.
Figures & tables
Coordinate
Physical quantity
Feasibility information
ccm
contact margin
remaining contact versus transmission time
ced
energy difference
selected offload versus local execution
ccv
coverage validity
binary reachability of the selected target
ccu
compute utilization
post-action load at the selected target
cdb
deadline buffer
deadline minus predicted completion time
Table 1: Coordinates of the verifiable feasibility representation. Every target is computed for the action already proposed by the actor.
Figure 1: VERA training and execution. The actor predicts VFR and conditions its action on both VFR and backbone features. The environment computes the exact target only after an action is proposed. The green verification and candidate-scoring path is removed at execution.
Protocol
Method
SR (%) ↑
COV (%) ↓
Standard S5
VERA
55.33 ± 3.60
0.45 ± 0.81
AB-MAPPO
50.68 ± 4.65
45.32 ± 9.42
MAPPO
26.35 ± 6.15
69.54 ± 11.07
Tuned constrained policy
40.17 ± 3.46
33.65 ± 8.67
MaskMAPPO †
56.66 ± 0.70
0.05
Reward-matched S10
VERA
56.00 ± 5.19
0.51 ± 0.71
Table 2: Primary SAGIN results. Values are mean ± sample standard deviation across seeds. Standard S5 and reward-matched S10 are separate protocols; MaskMAPPO † observes exact decision-time feasibility.
Figure 2: Mechanism and scope. (a) CGRA improves S10 SAGIN success across four representations. (b) VERA transfers across seven topologies; intervals use seed-level 95% confidence intervals from 20 episodes per seed. (c) All CGRA-bearing force/contact arms meet the pre-registered collapse rule. (d) Three stabilizers preserve SAGIN but pass neither cross-domain rescue gate. Standard-deviation bars are used in (a).
Figure 3: Scale, zero-shot shifts, and execution cost. (a) Mean SR as the number of users grows. (b) VERA SR retention under dimension-preserving shifts; the tighter deadline and heavier tasks also reduce intrinsic feasibility. (c) Per-step actor latency, reported as mean ± sample standard deviation. Panel titles are placed below each subplot.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Method
SR (%)
Latency (s)
Energy
Cov. viol. (%)
Jain
VERA
55.33 ± 3.60
0.1010 ± 0.0003
0.018 ± 0.015
0.45 ± 0.81
0.9936 ± 0.0007
AB-MAPPO
50.68 ± 4.65
0.0956 ± 0.0028
8.482 ± 2.689
45.32 ± 9.42
0.9825 ± 0.0092
B-MAPPO
49.14 ± 5.39
0.0950 ± 0.0040
11.849 ± 10.424
46.65 ± 10.16
0.9841 ± 0.0065
MAPPO
26.35 ± 6.15
0.0874 ± 0.0027
55.851 ± 40.247
69.54 ± 11.07
0.9779 ± 0.0057
MADDPG
21.10 ± 6.63
0.0722 ± 0.0076
6.219 ± 1.951
91.74 ± 10.19
0.8893 ± 0.0335
Random
14.86 ± 1.51
0.0829 ± 0.0025
179.987 ± 13.650
97.17 ± 1.95
0.9447 ± 0.0154
Appendix
Table 3: Standard configuration, final-500-episode window, five seeds. Latency is averaged over successful tasks and must be read with success rate.
Figure 4: Standard-configuration success across the complete method comparison. MaskMAPPO uses a privileged hard feasibility mask and is shown as a reference rather than a deployable competitor; the constrained baseline is the tuned configuration from Table 2 . Error bars are sample standard deviations across five seeds.
Variant
SR (%)
Cov. viol. (%)
VFR validity (%)
Full
54.57 ± 2.26
2.33 ± 1.44
82.63 ± 2.73
w/o VFR
51.98 ± 3.00
12.65 ± 9.03
–
w/o coverage coordinate
53.45 ± 0.51
0.78 ± 0.65
87.39 ± 0.83
w/o verification reward
37.10 ± 6.07
60.67 ± 13.16
54.93 ± 5.35
Appendix
Table 4: Five-seed ablation under the VERA execution objective.
Figure 5: Same-objective ablation of success, coverage violation, and representation validity. The larger reward-matched ladder and S10 factorial provide the stronger mechanism tests.
Method
SR (%)
Cov. viol. (%)
VERA
56.00 ± 5.19
0.51 ± 0.71
AB-MAPPO-R
47.26 ± 6.34
52.70 ± 10.55
Appendix
Table 5: Reward-matched SAGIN comparison over ten paired seeds.
Variant
VFR supervision
VFR enters action
Consistency
CGRA
Backbone
no
no
no
no
Latent-5D
no
yes
no
yes
Aux-VFR
yes
no
no
yes
VFR-cond.
yes
yes
no
yes
w/o CGRA
yes
yes
yes
no
Strict bottleneck
yes
yes, no backbone bypass
yes
no
Appendix
Table 6: Component switches in the reward-matched mechanism ladder.
Variant
SR (%) ↑
Cov. viol. (%) ↓
VFR valid. (%) ↑
Backbone
36.79 ± 2.33
62.50 ± 6.27
–
Latent-5D
45.89 ± 8.14
25.56 ± 17.55
17.60 ± 10.98
Aux-VFR
52.66 ± 0.62
2.40 ± 1.74
85.51 ± 4.03
VFR-cond.
54.65 ± 4.35
2.73 ± 4.40
79.32 ± 3.52
w/o CGRA
40.12 ± 5.42
62.23 ± 8.99
54.67 ± 6.24
Strict bottleneck
42.25 ± 1.81
52.05 ± 6.94
49.75 ± 4.02
Appendix
Table 7: Reward-matched mechanism ladder, five seeds. Aux-VFR is supervised but not fed to the action head; VFR-cond. adds that pathway.
Figure 6: Seven-way reward-matched mechanism ladder. (a) Converged success. (b) Coverage violation. Error bars are sample standard deviations across five seeds; panel titles appear below the subplots.
Variant
SR (%)
Δ SR (pp)
Cov. viol. (%)
Full
54.25 ± 1.29
0.00
2.78 ± 0.68
Drop coverage margin
53.10 ± 0.60
-1.15
2.46 ± 0.71
Drop energy
53.21 ± 0.94
-1.04
2.97 ± 0.70
Drop coverage-valid
53.22 ± 0.76
-1.03
2.67 ± 1.09
Drop compute
53.22 ± 0.37
-1.03
2.54 ± 0.43
Drop delay
52.51 ± 0.26
-1.74
2.50 ± 0.87
Appendix
Table 8: Five-seed semantic-coordinate controls. Small changes relative to Full rule out a claim that any single named coordinate is the primary performance driver.
Figure 7: Success under coordinate removal, random-target, and coordinate-permutation controls. Error bars are sample standard deviations over five seeds.
Representation
CGRA SR
no-CGRA SR
Δ SR
CGRA COV
no-CGRA COV
Handcrafted
53.36 ± 0.79
38.45 ± 3.08
+14.91
2.37 ± 0.63
62.51 ± 5.27
Learned
53.31 ± 0.52
37.25 ± 2.77
+16.06
2.32 ± 0.70
64.45 ± 5.66
Random
52.53 ± 0.32
36.06 ± 2.89
+16.48
2.73 ± 0.96
59.60 ± 9.56
Latent-5D
49.58 ± 4.28
35.42 ± 1.99
+14.16
19.32 ± 12.07
61.40 ± 2.89
Appendix
Table 9: S10 representation–CGRA factorial in SAGIN. COV is coverage violation.
Topology
VERA
AB-MAPPO-R
MAPPO-R
Regional
53.17 ± 2.84
46.01 ± 7.59
26.12 ± 3.22
Regional deterministic
53.05 ± 0.21
50.00 ± 3.42
23.03 ± 3.07
Walker
50.93 ± 2.78
42.34 ± 5.13
26.59 ± 3.04
Walker-star
51.42 ± 2.94
40.68 ± 6.08
26.29 ± 2.90
Phase shift
52.44 ± 1.92
49.65 ± 16.93
25.73 ± 7.02
Inclination shift
52.81 ± 1.97
41.52 ± 5.59
26.58 ± 2.70
Appendix
Table 10: Absolute success rate (%) over seven topologies.
Bin
Margin mean
VFR error
Coverage flip (%)
Violation (%)
0
-0.740
0.115
8.5
100.0
1
-0.456
0.068
0.3
100.0
2
-0.219
0.071
0.3
100.0
3
-0.025
0.069
0.4
63.8
4
0.133
0.063
0.0
0.0
5
0.270
0.057
0.0
0.0
Appendix
Table 11: Signed-margin calibration aggregated over the source topology.
Figure 8: Source-topology signed-margin calibration. The violation transition occurs around the feasibility boundary; the preceding table verifies the same ordering on four shifted topologies.
Variant
Goal success (%)
Training collision events (%)
Full
24.75 ± 4.09
0.64 ± 0.92
Latent-5D
28.41 ± 3.69
0.26 ± 0.37
No VFR / no CGRA
4.86 ± 4.20
0.51 ± 0.59
Reward-only constrained
1.56 ± 2.96
0.73 ± 1.02
Appendix
Table 12: VMAS in-domain performance over five training seeds.
Figure 9: VMAS in-domain performance over five training seeds: (a) goal success and (b) training collision events. In-domain collisions stay low for every variant, so the safety difference is driven by the out-of-distribution audit. The latent representation exceeds the hand-specified semantic head, which is retained as a negative result for universal semantic superiority.
Variant
Density goal
Density collision
Scale goal
Scale collision
Full
36.67 ± 13.94
1.93 ± 3.25
20.00 ± 32.60
4.10 ± 3.54
No VFR/CGRA
20.00 ± 21.73
2.60 ± 4.56
10.00 ± 13.69
0.00 ± 0.00
Latent-5D
53.33 ± 21.73
6.60 ± 11.20
50.00 ± 39.53
7.50 ± 6.55
Constrained
23.33 ± 22.36
7.80 ± 7.15
20.00 ± 27.39
1.60 ± 3.05
Appendix
Table 13: VMAS OOD audit, mean ± sample standard deviation across five inference seeds.
Representation
CGRA SR
no-CGRA SR
CGRA COV
no-CGRA COV
Handcrafted
63.47 ± 0.06
49.73 ± 1.25
100.00 ± 0.00
5.70 ± 11.59
Learned
63.47 ± 0.06
49.57 ± 0.52
100.00 ± 0.00
2.65 ± 4.31
Random
63.47 ± 0.06
49.29 ± 0.71
100.00 ± 0.00
1.43 ± 0.63
Latent-5D
63.47 ± 0.06
49.27 ± 0.63
100.00 ± 0.00
2.07 ± 1.47
Appendix
Table 14: S10 resource-allocation factorial.
Arm
Collapsed seeds
Converged SR (%)
Converged COV (%)
No VFR/CGRA
2/10
47.42 ± 13.52
26.59
Full, no CGRA
3/10
48.11 ± 10.12
29.32
CGRA only
10/10
2.44 ± 1.65
4.63
Full
10/10
0.87 ± 1.68
4.07
Full, no annealing
10/10
2.33 ± 3.51
4.33
Appendix
Table 15: Force/contact collapse attribution over ten seeds per arm.
Figure 10: Measured transfer boundary. S10 resource allocation gains success by violating shared capacity across all four representations (a,b). In the 4,000-row VMAS OOD audit, Latent-5D has the highest goal reach but also the highest normalized collision rate (c,d). Error bars are sample standard deviations across seeds.
Arm
SR (%)
Cov. viol. (%)
Target flip vs. Full (%)
Full
61.98
0.73
0.0
Backbone-only
62.43
0.34
3.5
VFR-only
48.58
36.67
39.2
Both-zero
35.56
40.33
45.6
Appendix
Table 16: Inference-pathway intervention results.
Figure 11: Inference-time pathway decomposition. The table additionally reports the both-zero arm; this diagnostic is interpreted jointly with the retrained ladder.
Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of agents, and additional requirements couple agents in ways that reward structure alone does not capture. We introduce Coordination Graphs for Constrained Multi-Agent Reinforcement Learning (CG-CMARL), a framework that addresses both challenges by combining coordination graphs with Lagrangian duality. The system decomposes the joint problem into pairwise regions, each served by a set of shared Q-functions, one for the primary objective and one for each of the constraints, so that the number of learned models is independent of the number of agents. At execution time, Max-Sum message passing coordinates actions across the factor graph, while a Lagrangian multiplier controls the objective--constraint tradeoff, allowing a single trained model to trace a Pareto front without retraining. We provide convergence guarantees under mild conditions, together with a compositional error bound that decomposes into separate interpretable sources, each traceable to a specific design choice and independently controllable. Experiments on cooperative navigation tasks (where teams of up to 10 agents must coordinate to reach target positions while satisfying pairwise constraints) show that our method produces Pareto fronts dominating established baselines trained at fixed reward-shaping ratios, while scaling to team sizes where centralized approaches become intractable.
Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders Jonsson
Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain
We present a distributed approach for constrained Multi-Agent Reinforcement Learning (MARL) that combines state-augmented policy learning with distributed consensus over dual variables. Our method targets systems where agents have separable dynamics but must coordinate to satisfy global resource constraints, a setting in which, as we demonstrate empirically, independent learning fails to produce feasible solutions because agents cannot determine appropriate individual contributions toward collective constraint satisfaction. The key technical contribution is showing that lightweight neighbor-to-neighbor consensus over Lagrange multipliers suffices for globally coordinated constraint enforcement while preserving the scalability of independent training. Each agent learns a single augmented policy offline, conditioned on both its local state and a dual variable encoding constraint feedback. During execution, agents reach agreement on this dual variable through local communication alone. We prove that under mild connectivity assumptions, the consensus error among agents' multipliers is bounded, and show that this translates to a bounded constraint violation that decreases with graph connectivity and the number of consensus rounds. Unlike centralized training with decentralized execution (CTDE) approaches, whose complexity grows at least quadratically with agent count, our method scales linearly in both training and execution. Experiments on smart grid demand response demonstrate that consensus coordination is \emph{essential for feasibility}: without it, agents satisfy grid capacity constraints only by indefinitely postponing demand, a degenerate non-solution. With consensus, agents converge to a shared dual variable and satisfy both grid constraints and demand fulfillment, scaling to thousands of agents while CTDE baselines are limited to dozens.
Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders Jonsson
Department of Engineering, University Pompeu Fabra
Modern agent frameworks compose planners, tool agents, remote services, and shared specialists into runtime delegation graphs, but their revocation APIs still resemble token or subtree invalidation. When one delegation is withdrawn, the runtime must know which agents lose authority while independently authorized agents keep working. We study this authority consistency problem and introduce VERA (Verifiable Edge Revocation for Agents), a verifier-checkable revocation contract and API emitted by agent-runtime adapters as signed evidence. Under disjunctive authority, revoking edge e invalidates exactly T_intent(e,G) = reach(G) \ reach(G \ {e}), the agents whose every authorizing root path used e. Used as a contract, this target exposes two runtime failures: tree cascades over-revoke shared agents, while deployer-scoped cascades under-revoke cross-domain descendants. In a LangGraph framework-replt cells repeated 20 times yield 500compiled-framework traces and 2,000 valid signed delegation decisions; 13/25 cells contain runtime multi-parsharing and 8/25 contain cross-deployer shies 500/500 target proofs, preserves all320 alternate-parent shared-agent cases that tree cascade revokes, and rejects unauthorized signers and omission attacks. Baseline replay over 1,9that holder/node and tree-style targetscannot express this behavior. We further validate schema portability on A2A, AutoGen, and CrewAI artifacts: nine traces, including five executable Cregned delegation events that pass schema and signature checks.