Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework
Authors: Jose Tupayachi, Xueping Li, Soham Das
Organizations: Oak Ridge National Laboratory Department of Industrial and Systems Engineering, University of Tennessee, Knoxville · Department of Industrial and Systems Engineering, University of Tennessee, Knoxville
The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.
Figures & tables
Figure 1: Terminal depletion: horizontal axis is the epoch offset within the final 2,000 epochs; vertical axis is c=1−BT/K ( T=H ). Panels (A)/(B) show IPPO/MAPPO. Colors denote budgets ε ; dotted horizontal lines mark their targets. Curves show seed means. Envelopes span the seed minimum–maximum. See Appendix F.1 for notation.
Figure 2: Harvest return: horizontal axis is the epoch offset within the final 2,000 epochs; vertical axis is undiscounted joint/team return (the trainer logs total harvest divided by K ). Panels (A)/(B) show IPPO/MAPPO; colors denote ε . Curves show seed means. Envelopes span their minimum–maximum. This is not the discounted objective JΣγ ; see Appendix F.1 .
Figure 3: Ecological price: horizontal axis is the epoch offset within the final 2,000 epochs; vertical axis is the shared PI multiplier λ . Panels (A)/(B) show IPPO/MAPPO; colors denote ε . Curves show seed means and envelopes span seed ranges. This λ is distinct from λGAE .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value
Agents; initial stock; carrying capacity
n=2 , B0=K=1000
Growth; catchability; episode horizon
r=0.3 , q=0.5 , H=60
Actions
Continuous effort ei∈[0,1]
Actor; critics
Squashed Gaussian; local (IPPO), central (MAPPO)
Learning objectives
Individual reward (IPPO); team reward (MAPPO)
Constraint feedback
Shared terminal-depletion PI multiplier
Appendix
Table A1: Design of the constrained MARL experiments. Each method–budget–seed combination trains a separate controller for 20,000 epochs. The horizon is H=60 for all experiments.
LLMs are increasingly deployed as agents that plan, use tools, and act over time. When they share persistent resources, such as compute pools or energy reserves, decisions by one agent affect the conditions faced by later agents. We study this coordination failure in a renewable energy commons. Four same-family GPT, Gemini, or Grok agents act in homogeneous self-play as electricity prosumers, instructed to maximize operational continuity. Holding aggregate residual demand and the decision protocol fixed, we vary the regeneration rate of a shared energy reserve from abundance to scarcity. All three families preserve the reserve when demand does not exceed peak renewable replacement, but over-appropriate it beyond that threshold (all nine exact scarcity contrasts survive Holm correction; largest adjusted p = 4.87e-5). The pattern is self-defeating: the same populations protect current service while undermining future service. At higher scarcity (rho = 1.2), early aggregate request pressure exceeds peak renewable replacement in every family and averages 1.21 times that level. Mean trajectories fall below the reserve level of maximum replenishment by rounds 5-7. Two offline benchmarks compare a social planner maximizing group-wide operational-service value with open access, where each prosumer maximizes its own value. At a discount factor of gamma = 0.95, both benchmarks sustain the reserve under the same dynamics. Realized depletion instead resembles outcomes under a more impatient open-access benchmark. The populations therefore behave like impatient optimizers at the level of the public trajectory. This system-level alignment failure would be missed by isolated-response evaluation.
Constrained Multi-agent reinforcement learning (CMARL) faces two intertwined challenges: the joint action space grows exponentially with the number of agents, and additional requirements couple agents in ways that reward structure alone does not capture. We introduce Coordination Graphs for Constrained Multi-Agent Reinforcement Learning (CG-CMARL), a framework that addresses both challenges by combining coordination graphs with Lagrangian duality. The system decomposes the joint problem into pairwise regions, each served by a set of shared Q-functions, one for the primary objective and one for each of the constraints, so that the number of learned models is independent of the number of agents. At execution time, Max-Sum message passing coordinates actions across the factor graph, while a Lagrangian multiplier controls the objective--constraint tradeoff, allowing a single trained model to trace a Pareto front without retraining. We provide convergence guarantees under mild conditions, together with a compositional error bound that decomposes into separate interpretable sources, each traceable to a specific design choice and independently controllable. Experiments on cooperative navigation tasks (where teams of up to 10 agents must coordinate to reach target positions while satisfying pairwise constraints) show that our method produces Pareto fronts dominating established baselines trained at fixed reward-shaping ratios, while scaling to team sizes where centralized approaches become intractable.
Santiago Amaya-Corredor, Miguel Calvo-Fullana, Anders Jonsson
Department of Engineering, Universitat Pompeu Fabra, Barcelona, Spain
In large-scale multi-agent systems with shared resource constraints, an upstream planner must iteratively evaluate candidate resource plans -- assessing feasibility, aggregate response, and marginal cost -- before committing to one. Lagrangian relaxation separates local decisions through a broadcast cost signal, but the planner still needs the cost-to-utilization response map to explore plan space, and this map depends on population composition that changes across planning cycles. We propose \emph{population-aware coordination interfaces}: learned primal and dual maps, conditioned on compact population summaries, that the planner queries inside its iterative loop. The primal map predicts aggregate utilization under a proposed cost trajectory; the dual map predicts the cost trajectory for a target plan. By encoding response-relevant population structure, these maps remain reliable across evolving populations without per-cycle retraining, and support coordination of large populations from compact subsamples. We additionally cast Sim2Real transfer as a backtestable procedure, enabling evaluation before deployment. In a supply-chain capacity-control case study, population-aware interfaces reduce forecast error by 16--19% and capacity violations by 20--51% relative to population-unaware baselines under composition shift; 20K-agent cohorts support accurate coordination of 500K-agent populations; and simulator-trained primal maps achieve 11.1% MAPE on real observations versus 13--24% for baselines.