When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multipliers may still rely on the same aggregate cost signal. We introduce Lagrangian Responsibility Allocation (LiRA), which learns each agent's share of a common multiplier by optimizing social welfare over a finite training horizon. The multiplier enforces the aggregate budget, while responsibility shares redistribute its influence without modifying the original rewards or constraints. For convex games under standard regularity conditions, varying these shares induces a smooth family of normalized generalized Nash equilibria in which active constraints remain at their budgets while welfare varies. To optimize responsibility before convergence, we derive a welfare gradient that accounts for both learning updates and the induced change in data distribution. Across CityLearn, MABIM, Harvest, and MetaDrive, spanning 3 to 400 agents, LiRA improves average social welfare by up to 29% over uniform and agent-specific multiplier baselines. Grid and driving costs remain within budget, inventory violations decrease, and Harvest makes more effective use of available budget.
Figures & tables
Figure 1: Overview of LiRA. (a) In the illustrated corridor scenario, uniform shares leave both robots waiting; learned shares let the robot without an urgent package yield so that the loaded robot proceeds. Both outcomes avoid collision but differ in welfare. (b) Under Theorem 1 ’s conditions, varying responsibility ρ selects a smooth equilibrium family. For each active constraint k , cost Ck stays at its budget dk , while welfare W can vary and the shared equilibrium multiplier λ∗ adapts to ρ . (c) From checkpoint x0 , LiRA runs M≥2 independent lookaheads of q learner updates, each followed by fresh evaluation. It combines direct unrolling (DU) with sampling correction (SC) to estimate ∇ϕVq , update ϕ , and restart from x0 .
Task (N,K)
Outcome
Budget
Uniform
PAL
LiRA
CityLearn (3,1)
Welfare ( 103 )
–
−16.87±3.04
−16.87±3.04
−14.51±1.66
Grid excess (kWh)
28.76
28.44±5.97
28.44±5.97
28.73±7.48
MABIM (400,2)
Welfare ( 106 )
–
−596.31±5.00
−597.27±3.76
−591.36±4.48
Rejections C1 ( 106 )
20.35
21.51±0.25
21.56±0.23
21.30±0.18
Rejections C2 ( 103 )
32.69
27.71±2.31
27.54±2.21
28.68±0.90
Harvest (7,1)
Welfare
–
42.73±12.22
44.73±3.79
50.27±2.53
Table 1: Welfare and shared costs across four tasks. Held-out mean ± sample standard deviation over matched training seeds (three per task; six for MetaDrive). Costs and budgets use the units shown in each row. Bold marks the highest mean welfare.
Figure 2: Hourly temperature relative to setpoint in CityLearn. Each point averages observations at that hour in the matched 719-step seed-1102 evaluation. Uniform is dashed; LiRA is solid. Gray marks the ±1∘ C comfort band, and tan the 16–18h price peak. The episode totals appear in Table 2 .
Learned ρ
Electricity change (kWh)
Reward change ( 103 )
Building 1
0.395
−37.6
+4.43
Building 2
0.387
−32.7
+0.48
Building 3
0.218
−18.9
−0.03
Table 2: CityLearn responsibility and episode outcomes. Seed-1102 evaluation over 719 steps. Uniform assigns ρ=1/3 to each building; changes are LiRA minus Uniform. Reward is native building comfort reward.
Varied choice
Setting
Welfare ( 103 )
Grid excess (kWh)
Main setting
(2,8,0.05)
−14.51±1.66
28.73±7.48
Lookahead
q=3
−18.13±4.42
29.34±4.67
q=4
−19.03±8.39
31.52±4.20
Replicates
M=4
−16.02±4.05
28.91±4.02
M=12
−16.68±3.87
29.12±2.38
Step size
ηρ=0.025
−20.67±7.14
32.97±5.65
Table 3: CityLearn local search around the main setting. Each row changes one parameter from (q,M,ηρ)=(2,8,0.05) . Welfare and grid excess are mean ± sample standard deviation over three matched seeds. The grid-excess budget is 28.76 kWh.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Cold excess
ΔRcold
ΔRwarm
ΔRin-band
( ∘ C ⋅ steps)
Building 1
586.1
+4528.01
−102.80
+6.55
Building 2
214.5
+813.09
−340.49
+6.26
Building 3
70.2
+159.58
−205.55
+11.62
Appendix
Table 4: CityLearn comfort components in the matched seed-1102 episode. Cold excess is measured under Uniform; reward changes are LiRA minus Uniform.