Self-Referenced Social Preferences: Cooperation without Observing Others Rewards
Organizations: Amazon Vancouver, Canada · Nvidia Santa Clara, United States · Itaú Unibanco São Paulo, Brazil
Abstract
Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.
Figures & tables
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| Learner | PPO reward | actor coefficient | target |
|---|---|---|---|
| Plain PPO | own GAE | – | |
| TR | own GAE | – | |
| SRR | own GAE | ||
| SRA | |||
| + Standing | above | unchanged | unchanged |
| Game | Observation | Actions |
|---|---|---|
| Escape Room | one-hot place (start, lever, door) of every agent, its own first; values; no partial view | 3: stay, lever, door |
| Clean Up | window centred on the agent, 7 binary planes: self, others, apple, waste, river, wall, beam; full, partial | 6: four moves, stay, clean |
| Commons Harvest | RGB window centred on the agent and turned to its facing direction, self and others each in one fixed colour; full, partial | 8: four moves, stay, two turns, fire |
| Escape Room | Clean Up | Commons Harvest | |
|---|---|---|---|
| Shared network | MLP, two layers of 64 tanh units | convolutions with 16 and 32 channels (second with stride 2), 128 ReLU units | CleanRL Atari network: convolutions /4 (32), /2 (64), /1 (64), 512 ReLU units |
| Reward model | MLP, two layers of 64 tanh units, on | two convolutions (16 channels) on the stacked , max- and mean-pooled, then 64 units with | one layer of 64 tanh units on the shared network’s features of and (not trained through) and |
| Steps; parallel envs | 500k; 16 | 5M; 8 | 5M; 8 |
| Rollout per env | 16, 32, or 128 (tuned) | 256 | 128 |
| Learning rate | or (tuned) | ||
| Entropy coefficient | 0.01 or 0.05 (tuned) | 0.01 ( ); 0.05 ( ) ∗ | 0.01 |
| Learner (weights) | LR, entropy | Rollout |
|---|---|---|
| Plain PPO | , 0.01 | 32 |
| TR-EI ( ) | , 0.01 | 128 |
| TR-IA ( ) | , 0.01 | 32 |
| TR-SVO ( , ) | , 0.01 | 128 |
| SRA-EI ( ) | , 0.05 | 128 |
| SRA-IA ( ) | , 0.05 | 128 |
| selected on collective return | selected on Nash welfare | |||||
| Learner | ||||||
| Escape Room ( ) | ||||||
| Plain PPO | same setting | |||||
| TR-EI | same setting | |||||
| SRA-EI | same setting | |||||
| SRR-EI | same setting | |||||
| Method (selection) | Sust. | |||
|---|---|---|---|---|
| Plain PPO | ||||
| TR-EI (both) | ||||
| TR-SVO (both) | ||||
| TR-IA (C) | ||||
| TR-IA (NW) | ||||
| SRA-EI (both) |
| Env. | Method (selection) | Behavior | |||
| Escape Room | Plain PPO | no reliable escape | |||
| TR-EI (C) | concentrated exits, SD 0.49 | ||||
| SRA-EI | concentrated exits, SD 0.48 | ||||
| SRR-IA + Standing (NW) | criterion 8/8, SD 0.13 | ||||
| Clean Up | Plain PPO | – | |||
| TR-SVO (C) | cleaning share |
| PPO | TR | Best self-ref. | SRA-EI | |||
|---|---|---|---|---|---|---|
| 2 | 1 | 0.00 | 1.00 | 1.00 ( SRR-EI ) | 1.00 | 7/16 |
| 3 | 2 | 0.00 | 1.00 | 1.00 ( SRR-IA ) | 0.69 | 3/16 |
| 5 | 3 | 0.01 | 1.00 | 1.00 ( SRR-IA ) | 0.93 | 3/16 |
| 7 | 4 | 0.00 | 1.00 | 1.00 ( SRR-IA ) | 0.82 | 2/16 |
| 10 | 5 | 0.00 | 1.00 | 1.00 ( SRR-IA ) | 0.69 | 2/16 |
| 12 | 6 | 0.00 | 1.00 | 1.00 ( SRR-IA ) | 0.00 | 2/16 |
| Method | Setting | Returns | |||
|---|---|---|---|---|---|
| Plain PPO | – | 172 / 153 / 48 | |||
| SRA-EI (NW) | 335 / 260 / 173 | ||||
| SRA-SVO | 369 / 242 / 196 | ||||
| SRA-IA | 257 / 223 / 209 | ||||
| SRR-IA+V | 296 / 207 / 160 | ||||
| TR-IA | 361 / 243 / 216 |
| Method | Setting | Returns, high low | |||
|---|---|---|---|---|---|
| Plain PPO | – | 70 / 60 / 51 / 49 / 47 / 31 / 5 | |||
| SRA-EI (NW) | 426 / 409 / 383 / 352 / 285 / 64 / 34 | ||||
| SRA-SVO (NW) | 266 / 252 / 244 / 239 / 231 / 85 / 39 | ||||
| SRR-IA (NW) | 178 / 166 / 162 / 155 / 150 / 142 / 116 | ||||
| SRR-IA+V (NW) | 181 / 166 / 154 / 143 / 133 / 129 / 100 | ||||
| TR-IA (NW) | 76 / 65 / 62 / 55 / 52 / 23 / 11 |
| ER( , ), with-standing configuration | Exit SD | Shared seeds | ||
|---|---|---|---|---|
| (3,2) SRA-EI , | ||||
| (5,3) SRA-EI , | ||||
| (5,3) SRR-IA , | ||||
| (5,3) TR-IA , † | ||||
| (7,4) SRR-IA , | ||||
| (7,4) TR-IA , † |
| Control | Gate | |||
|---|---|---|---|---|
| Critic trained on raw returns | off | |||
| on | – | |||
| Standing as a policy input | off | |||
| on | – | |||
| One shared true ledger | on | – | ||
| with standing input | on | – |
| Method | Rewards | Obs./act. | Policies | Critic |
|---|---|---|---|---|
| Inequity aversion ( Hughes et al., 2018 ) | ||||
| Social value orientation ( McKee et al., 2020 ) | ||||
| Prosocial reward ( Peysakhovich and Lerer, 2018 ) | ||||
| VDN, QMIX ( Sunehag et al., 2018 ; Rashid et al., 2018 ) | (team) | |||
| COMA ( Foerster et al., 2018b ) | (team) | |||
| MADDPG ( Lowe et al., 2017 ) |