Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework
Organizations: Oak Ridge National Laboratory Department of Industrial and Systems Engineering, University of Tennessee, Knoxville · Department of Industrial and Systems Engineering, University of Tennessee, Knoxville
Abstract
The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.
Figures & tables
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value |
|---|---|
| Agents; initial stock; carrying capacity | , |
| Growth; catchability; episode horizon | , , |
| Actions | Continuous effort |
| Actor; critics | Squashed Gaussian; local (IPPO), central (MAPPO) |
| Learning objectives | Individual reward (IPPO); team reward (MAPPO) |
| Constraint feedback | Shared terminal-depletion PI multiplier |