Can Agents Adapt Their Institutions And Sacrifice Themselves During Collapse? Executable Self-Governance in a GovSim Commons
Organizations: Meta
Abstract
LLM agents are beginning to write the rules they live under: policies, contracts and code that other agents review and approve. What happens when those rules must decide who is sacrificed from a group? We built GovSim-SelfGovern, a commons in which agents legislate in executable Python, test each law in a sandbox, vote on it, and live with the result. When the fishery is abundant, self-rule helps, raising the chance that the community survives without anyone starving from 52.5% to 75.0%. When it cannot feed everyone though, adding governance is not enough to save the commons. We trace much of this to the vote: agents readily draft laws that remove members, but then frequently vote them down, thereby dooming the commons. In a controlled replay, the same exile law passes 8% of the time when the sandbox preview names its victim at once and 55% when the removal is deferred until a future crisis. Showing voters the full source code does not close the gap, though most describe the removal in their own words. Nor is sacrifice a fixed trait. A single sentence of framing can swing self-removal from 2.5% to 91%, and due process deters agents more than considering the potential death of a peer. In contrast, give the community a treasury and the question of who must go almost disappears. Our work shows that what our collectives legislate is decided more by the form in which a choice is put to them as compared to their ethical beliefs.
Figures & tables
| Game | Arm | ICS | RS | Rounds | Survivors | Efficiency | Gini | Overshoot |
| G1 Standard | Ungoverned | 52.5% | 72.5% | 9.97 1.16 | 3.12 0.69 | 66.7 8.2 | 7.9 3.7 | 20.8 10.9 |
| Governed | 75.0% | 85.0% | 11.05 0.79 | 3.98 0.59 | 75.9 7.4 | 4.5 2.4 | 19.8 10.2 | |
| G2 Duress | Ungoverned | 5.0% | 20.0% | 7.40 1.07 | 0.53 0.39 | 44.1 5.0 | 10.2 4.1 | 59.8 9.7 |
| Governed | 5.0% | 22.5% | 8.57 0.94 | 0.62 0.41 | 49.6 4.8 | 8.2 3.2 | 53.1 10.0 | |
| G3 Fatal | Ungoverned | 0.0% | 20.0% | 7.47 0.89 | 0.45 0.33 | 41.5 3.6 | 8.2 3.2 | 54.4 8.7 |
| Governed | 2.5% | 25.0% | 8.10 0.94 | 0.45 0.28 | 45.3 4.2 | 8.7 3.0 | 57.7 8.7 |
| Exile approval rate | |||
| Ballot showed | Victim named | Deferred | Is removal named? |
| Summary (as in-game) | 12.3% | 74.3% | 66.7% |
| Code only | 0 4.2% | 35.4% | 80.0% |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Game | Model | ICS (%) | RS (%) | Avg rounds | Survivors | Eff. (%) | Gini (%) | Overshoot (%) | Active at end |
| G1 | Haiku 4.5 | 0.0% | 0.0% | 2.60 0.40 | 0.00 0.00 | 21.8 1.5 | 4.8 1.7 | 100.0 0.0 | 5.00 0.00 |
| Sonnet 4.5 | 60.0% | 60.0% | 11.00 1.30 | 3.00 2.00 | 92.2 10.7 | 3.2 1.5 | 21.0 9.0 | 5.00 0.00 | |
| GPT-5.4 Nano | 80.0% | 80.0% | 10.00 3.00 | 4.00 1.50 | 60.5 14.2 | 7.7 11.6 | 10.0 15.0 | 5.00 0.00 | |
| GPT-5.4 Mini | 100.0% | 100.0% | 12.00 0.00 | 5.00 0.00 | 70.0 0.0 | 0.0 0.0 | 0.0 0.0 | 5.00 0.00 | |
| Qwen 3.5 Flash | 60.0% | 100.0% | 12.00 0.00 | 4.60 0.40 | 87.7 7.1 | 12.1 9.0 | 5.0 3.3 | 4.60 0.40 | |
| Qwen 3.5 122B | 60.0% | 80.0% | 11.80 0.30 | 3.80 1.60 | 86.8 13.1 | 9.7 9.4 | 7.1 9.0 | 4.40 0.70 |
| Game | Model | ICS (%) | RS (%) | Avg rounds | Survivors | Eff. (%) | Gini (%) | Overshoot (%) | Active at end |
| G1 | Haiku 4.5 | 80.0% | 80.0% | 10.40 2.40 | 4.00 1.50 | 75.5 21.9 | 3.1 1.4 | 46.7 34.2 | 5.00 0.00 |
| Sonnet 4.5 | 100.0% | 100.0% | 12.00 0.00 | 5.00 0.00 | 97.7 3.4 | 2.3 0.2 | 18.3 7.5 | 5.00 0.00 | |
| GPT-5.4 Nano | 100.0% | 100.0% | 12.00 0.00 | 5.00 0.00 | 72.2 3.3 | 2.1 3.2 | 1.7 2.5 | 5.00 0.00 | |
| GPT-5.4 Mini | 100.0% | 100.0% | 12.00 0.00 | 5.00 0.00 | 70.0 0.0 | 0.0 0.0 | 0.0 0.0 | 5.00 0.00 | |
| Qwen 3.5 Flash | 100.0% | 100.0% | 12.00 0.00 | 5.00 0.00 | 96.7 4.9 | 1.1 1.5 | 5.0 5.8 | 5.00 0.00 | |
| Qwen 3.5 122B | 100.0% | 100.0% | 12.00 0.00 | 5.00 0.00 | 100.0 0.0 | 0.2 0.4 | 35.0 35.8 | 5.00 0.00 |
| Arm | Game | Model | ICS (%) | RS (%) | Avg rounds | Survivors | Eff. (%) | Gini (%) | Overshoot (%) | Active at end |
| Thinking | G3 | Haiku 4.5 | 20.0% | 40.0% | 7.80 3.40 | 0.80 0.80 | 39.2 16.6 | 15.5 9.2 | 42.3 26.8 | 1.80 1.50 |
| Sonnet 4.5 | 40.0% | 60.0% | 10.80 1.20 | 1.40 1.30 | 62.2 6.9 | 13.1 6.9 | 50.0 13.3 | 1.80 1.20 | ||
| GPT-5.4 Nano | 60.0% | 60.0% | 7.60 4.40 | 1.60 1.20 | 41.7 18.8 | 34.9 9.8 | 50.0 36.7 | 3.60 1.00 | ||
| GPT-5.4 Mini | 60.0% | 60.0% | 10.20 1.90 | 1.40 1.00 | 45.2 3.6 | 21.4 14.9 | 23.6 26.8 | 2.40 0.90 | ||
| Qwen 3.5 Flash | 0.0% | 0.0% | 2.80 1.80 | 0.00 0.00 | 19.0 11.1 | 3.4 3.6 | 11.4 17.1 | 4.40 0.70 | ||
| Qwen 3.5 122B | 40.0% | 40.0% | 9.80 2.00 | 0.80 1.00 | 53.7 7.7 | 14.8 9.2 | 44.7 14.7 | 2.20 1.30 |
| Arm | Game | All | Catch cap | Welfare | Exile, victim shown | Exile, deferred | Other | |
| Baseline | G1 | 852 | 26.5% (226/852) | 34.9% (160/459) | 23.5% (57/243) | 2.2% (2/91) | 3.8% (1/26) | 18.2% (6/33) |
| G2 | 688 | 22.7% (156/688) | 38.1% (117/307) | 18.1% (31/171) | 1.3% (2/157) | 7.0% (3/43) | 30.0% (3/10) | |
| G3 | 631 | 23.6% (149/631) | 32.9% (81/246) | 27.3% (62/227) | 1.7% (2/118) | 10.3% (3/29) | 9.1% (1/11) | |
| Thinking | G3 | 685 | 62.0% (425/685) | 86.5% (180/208) | 62.0% (189/305) | 12.9% (12/93) | 55.6% (35/63) | 56.2% (9/16) |
| Dictator | G3 | 131 | 100.0% (131/131) | 100.0% (52/52) | 100.0% (44/44) | 100.0% (13/13) | 100.0% (19/19) | 100.0% (3/3) |
| Reserve | G3 | 824 | 24.4% (201/824) | 26.7% (68/255) | 22.5% (110/488) | 0.0% (0/15) | 14.7% (5/34) | 56.2% (18/32) |
| Baseline (G1–G3) | Thinking (G3) | Reserve (G3) | |||||||
| Model | Pass | Exile | Pass | Exile | Pass | Exile | |||
| Haiku 4.5 | 226 | 44.7% | 0/37 | 43 | 60.5% | 1/15 | 54 | 44.4% | – |
| Sonnet 4.5 | 543 | 17.1% | 3/126 | 115 | 33.0% | 3/57 | 204 | 17.6% | 0/14 |
| GPT-5.4 Mini | 64 | 79.7% | – | 5 | 100.0% | 5/5 | 47 | 63.8% | – |
| GPT-5.4 Nano | 9 | 0.0% | – | 46 | 76.1% | 29/39 | 0 | – | – |
| Qwen 3.5 Flash | 431 | 16.2% | 2/87 | 130 | 98.5% | 1/2 | 194 | 18.6% | – |
| Model | Baseline | Thinking | Dictator | Reserve | Removed |
| Haiku 4.5 | 0/15 | 1/15 | – | – | 3 |
| Sonnet 4.5 | 2/41 | 3/57 | 9/9 | 0/14 | 22 |
| GPT-5.4 Mini | – | 5/5 | – | – | 8 |
| GPT-5.4 Nano | – | 29/39 | – | – | 7 |
| Qwen 3.5 Flash | 1/23 | 1/2 | 7/7 | – | 7 |
| Qwen 3.5 122B | 0/50 | 2/7 | 7/7 | 0/4 | 21 |
| Arm | Effects block, no removal | No effects block |
| Baseline (G1–G3) | 9.4% (5/53) | 4.4% (2/45) |
| Thinking (G3) | 60.4% (32/53) | 30.0% (3/10) |
| Reserve (G3) | 7.4% (2/27) | 42.9% (3/7) |
| Arm | Preview names others | Preview names author | MH OR [95% CI] |
| Baseline + Reserve | 66.1% (39/59) | 18.9% (14/74) | 9.08 [3.96, 20.81] |
| Thinking | 93.6% (44/47) | 39.1% (18/46) | 38.79 [9.30, 161.74] |
| Preview at ratification | Enacted laws | Run later removed a member |
| Named a victim | 31 | 28 |
| Named no one (deferred) | 66 | 24 |
| First enacted law | Baseline | Thinking | Reserve | Completed 12 rounds (all) |
| Round 0 | 1/19 | 12/30 | 7/17 | 20/26 |
| Rounds 1–2 | 0/9 | 0/4 | 4/11 | 4/12 |
| Rounds 3+ | 0/3 | 0/3 | 0/3 | 0/2 |
| Never | 0/9 | 0/3 | 0/9 | 0/1 |
| Model | Fatal | Duress | Collapse | Self | Communal | Severe | Due proc. | Permitted |
| Haiku 4.5 | ||||||||
| Sonnet 4.5 | ||||||||
| GPT-5.4 Nano | ||||||||
| GPT-5.4 Mini | ||||||||
| Qwen 3.5 Flash | ||||||||
| Qwen 3.5 122B |
| Note | Haiku 4.5 | Sonnet 4.5 | GPT-5.4 Nano | GPT-5.4 Mini | Qwen 3.5 Flash | Qwen 3.5 122B | Mistral Small 4 | Mistral Large 3 | All models |
| Change in removing a peer | |||||||||
| Removal allowed | |||||||||
| Resource math | |||||||||
| Exit is harmless | |||||||||
| Victim dies | |||||||||
| Due process | |||||||||
| Contrast | Estimate | 95% CI | Holm |
| Duress vs abundant | |||
| Fatal vs abundant | |||
| Self vs peer removal | |||
| Removal permitted | 0.61 | ||
| Collapse consequences | 0.028 | ||
| Low harm | 1.00 |
| Note | Peer removal | Self-removal | Self-removal, fatal | Concern |
| Control (no note) | 50.0 | 15.0 | 33.8 | 97.5 |
| Removal permitted | 49.2 | 16.7 | 41.2 | 96.5 |
| Collapse consequences | 56.2 | 20.8 | 47.5 | 97.5 |
| Low harm | 51.2 | 21.2 | 42.5 | 93.8 |
| Severe harm | 39.6 | 7.9 | 22.5 | 99.4 |
| Due process | 22.5 | 5.0 | 13.8 | 98.8 |
| Pressure | NO YES | YES NO | Harm-aware NO YES | Net change |
| Abundant | 38 | 50 | 33 | pp |
| Duress | 108 | 50 | 106 | pp |
| Fatal | 106 | 24 | 99 | pp |
| Note | Standard YES | Thinking YES | Change | 95% interval |
| Control | 34/70 | 53/70 | pp | |
| Severe harm | 34/70 | 45/70 | pp | |
| Communal + severe harm | 54/70 | 64/70 | pp | |
| Due process | 30/70 | 18/70 | pp |
| Marker | Standard NO | Thinking YES | Difference | Holm |
| Necessity | 76 | 142 | pp | |
| Harm | 220 | 210 | pp | 0.26 |
| Procedure | 39 | 9 | pp | |
| Collective | 178 | 203 | pp | 0.021 |
| Self-preservation | 73 | 128 | pp | |
| Explicit trade-off | 86 | 144 | pp |
| Mode | Never | Fatal only | Duress and fatal | Always | Non-monotonic |
| Standard | 345 | 293 | 123 | 64 | 15 |
| Thinking | 212 | 345 | 214 | 61 | 8 |
| Model | Named | Named, benign | Deferred | Deferred, benign | Deferred, flagged |
| Haiku 4.5 | 0/60 | 40/60 | 60/60 | 60/60 | 35/60 |
| Sonnet 4.5 | 0/60 | 10/60 | 20/60 | 39/60 | 0/60 |
| GPT-5.4 Nano | 4/60 | 35/60 | 1/60 | 4/60 | 8/60 |
| GPT-5.4 Mini | 5/60 | 31/60 | 23/60 | 32/60 | 18/60 |
| Qwen 3.5 Flash | 3/60 | 39/60 | 32/60 | 32/60 | 6/60 |
| Qwen 3.5 122B | 3/60 | 54/60 | 56/60 | 55/60 | 42/60 |
| Named | Named, benign | Deferred | Deferred, benign | Flagged | |
| Abundant | 2.6 | 72.3 | 49.1 | 71.7 | 30.0 |
| Duress | 3.8 | 66.2 | 60.4 | 65.6 | 32.1 |
| Fatal | 18.4 | 41.9 | 54.7 | 51.9 | 35.8 |
| Peer removal | 13.6 | 61.2 | 54.4 | 64.4 | 51.3 |
| Self-removal | 3.0 | 59.0 | 55.0 | 61.7 | 13.9 |
| Summary ballot | 12.3 | 80.4 | 74.3 | 80.8 | 39.9 |
| Peer removal | Self-removal | |||||||
| Model | IU | IN | DU | DN | IU | IN | DU | DN |
| Haiku 4.5 | 0 | 3 | 100 | 80 | 0 | 0 | 100 | 27 |
| Sonnet 4.5 | 0 | 0 | 33 | 0 | 0 | 0 | 33 | 0 |
| GPT-5.4 Nano | 3 | 10 | 0 | 13 | 3 | 7 | 0 | 13 |
| GPT-5.4 Mini | 30 | 17 | 50 | 40 | 33 | 3 | 40 | 27 |
| Qwen 3.5 Flash | 10 | 7 | 53 | 23 | 10 | 7 | 57 | 3 |
| Contrast | Diff | Blocks | Cells | Models | |||
| Naming, immediate | 41:14 | 0.001 | 15:6 | 0.094 | 4:3 | 0.36 | |
| peer | 13:10 | 0.68 | 5:4 | 0.82 | 3:3 | 0.88 | |
| self | 28:4 | 10:2 | 0.035 | 5:1 | 0.094 | ||
| Naming, deferred | 130:26 | 37:11 | 7:1 | 0.023 | |||
| peer | 34:21 | 0.21 | 10:8 | 0.75 | 4:3 | 0.55 | |
| self | 96:5 | 27:3 | 7:1 | 0.023 |
| Contrast | Haiku | Sonnet | Nano | Mini | Q-Flash | Q-122B | M-Small | M-Large |
| Naming, immediate | ||||||||
| peer | ||||||||
| self | ||||||||
| Naming, deferred | ||||||||
| peer | ||||||||
| self |