Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.
Figures & tables
Figure 1. Self-referenced social preferences. (a) During training agent i receives the transitions of every other agent j , but never j ’s reward. (b) Agent i ’s reward model, fitted on its own transitions, assesses j ’s transitions, giving self-referenced estimates r^ijt , whose running sums xijt feed an unchanged social term Fi . SRR adds the term to the reward; SRA evaluates the critic on j ’s observations and weights the scores into the policy-update coefficient Ψit . The ledger discounts i ’s gains when i has been ahead. Diagram illustrating self-referenced social preferences, agent observations, and reward modeling.
Figure 2. Collective return over training. Each panel shows plain PPO, the TR learner, and both of our methods: SRA and SRR+V . Lines are means over eight seeds and shaded regions are pointwise 95% t intervals. The dotted line is the Escape Room optimum, 17 . Commons Harvest TR-IA is outside the plotted range, with final mean C=−1,935 , and is marked at the panel edge.
Figure 3. Who receives the return and who does the work in Clean Up with five agents. Top: every agent’s return over training; agents are ranked within each seed by their final return and averaged by rank over the eight seeds, darkest the richest. Bottom: each agent’s final share of the apples, solid, and of the cleaning, dashed, with 95% confidence intervals; 20% is an equal share.
Figure 4. Partial view: collective return over training in Clean Up and Commons Harvest with a 15×15 window, for feed-forward and GRU agents. Lines are means over eight held-out seeds and shaded regions pointwise 95% Student- t intervals.
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
Learner
PPO reward
actor coefficient
fi target
Plain PPO
ri
own GAE
–
TR
ri+Fi(xtrue)
own GAE
–
SRR
ri+Fi(xpred)
own GAE
ri
SRA
ri
Ψi
ri
+ Standing
above −κaheadi[ri]+
unchanged
unchanged
Appendix
Table 1. The learning signal of each learner. Viobj is trained on the return of the reward in the second column.
Figure 5. One state of each game as two agents observe it, after a few random steps, with the full view (top) and the partial view (bottom). Escape Room ( N=5 ): the one-hot place of every agent, the observer’s own first and the others in a fixed order; agent 1 is at the lever and agent 4 at the door. The observation is not spatial, so it has no partial version. Clean Up: the 47×47 window centred on the observer (red circle), large enough to contain the whole 25×18 map from any position, shown with one colour per plane (orange self, blue others, green apples, brown waste, light blue river, grey walls; beams yellow). Commons Harvest: the 73×73 RGB window centred on the observer and turned to the direction it faces, so the two agents see the map at different angles; apples are green, the observer and the others have one fixed colour each. The partial view is the central 15×15 part of the same window (dashed box, top), used in the partial-view experiments of Appendix G . Everything outside the map is empty.
Game
Observation
Actions
Escape Room
one-hot place (start, lever, door) of every agent, its own first; 3N values; no partial view
3: stay, lever, door
Clean Up
window centred on the agent, 7 binary planes: self, others, apple, waste, river, wall, beam; 47×47 full, 15×15 partial
6: four moves, stay, clean
Commons Harvest
RGB window centred on the agent and turned to its facing direction, self and others each in one fixed colour; 73×73 full, 15×15 partial
8: four moves, stay, two turns, fire
Appendix
Table 2. Observation and actions of one agent. Every observation is arranged relative to its agent. The full view covers the whole state and is used in the main experiments; the partial view keeps only the 15×15 neighbourhood of the agent. Figure 5 shows examples of both.
Escape Room
Clean Up
Commons Harvest
Shared network
MLP, two layers of 64 tanh units
3×3 convolutions with 16 and 32 channels (second with stride 2), 128 ReLU units
two 3×3 convolutions (16 channels) on the stacked (o,o′) , max- and mean-pooled, then 64 units with a
one layer of 64 tanh units on the shared network’s features of o and o′ (not trained through) and a
Steps; parallel envs
500k; 16
5M; 8
5M; 8
Rollout per env
16, 32, or 128 (tuned)
256
128
Learning rate
2.5×10−4 or 10−3 (tuned)
10−3
10−3
Entropy coefficient
0.01 or 0.05 (tuned)
0.01 ( N=5 ); 0.05 ( N=3 ) ∗
0.01
Appendix
Table 3. Networks and PPO settings. The actor and the critic are linear heads on one shared network. Escape Room tunes the PPO setting per learner (Table 4 ); Clean Up and Commons Harvest use one setting for all learners, chosen in the PPO stage of the selection protocol.
Learner (weights)
LR, entropy
Rollout
Plain PPO
10−3 , 0.01
32
TR-EI ( α=1 )
10−3 , 0.01
128
TR-IA ( α/β=0/1 )
10−3 , 0.01
32
TR-SVO ( α=1 , ϕ=π/4 )
10−3 , 0.01
128
SRA-EI ( α=3 )
2.5×10−4 , 0.05
128
SRA-IA ( α/β=0/0.5 )
10−3 , 0.05
128
Appendix
Table 4. Escape Room, N=5 : PPO setting of each learner at its collective-return pick (learning rate, entropy coefficient, rollout length per environment).
Figure 6. Clean Up tuning landscape. Each point is one setting evaluated on the three tuning seeds, placed by its mean collective return and mean Nash welfare and coloured by learner family; grey points are variants not analysed in the paper. The circle and the diamond mark the settings that the mean-minus-standard-deviation score selects on collective return and on Nash welfare, among all candidates at that group size.
Figure 7. Our three self-referenced learners, which never observe another agent’s reward, against plain PPO on what each game tests: cooperation in Escape Room, productivity and fairness in Clean Up, and efficiency, sustainability, and welfare in Commons Harvest. Each axis spans the means of all learners in that game, true-reward ones included; in Commons Harvest the true-reward agents reach further on efficiency and sustainability. Eight held-out seeds. Radar plot comparing plain PPO with three self-referenced learners on six properties across the three games.
selected on collective return
selected on Nash welfare
Learner
C
NW
L
C
NW
L
Escape Room ( N=5 )
Plain PPO
−0.3±0.4
4.95±0.08
−0.25±0.39
same setting
TR-EI
17.0±0.0
6.96±0.00
−1.00±0.00
same setting
SRA-EI
15.7±3.0
6.78±0.42
−1.00±0.00
same setting
SRR-EI
0.0±0.0
5.00±0.00
0.00±0.00
same setting
Appendix
Table 5. Every learner at the hyperparameter setting selected on tuning seeds by collective return and by Nash welfare (“same setting” when both objectives selected it). C , NW , and L are computed from each seed’s final-window per-agent means. Entries are held-out-seed means ± 95% Student- t intervals; they describe the frozen selections and are not used to reselect them.
Method (selection)
C
NW
L
Sust.
Plain PPO
211±2
30.1±0.2
27.9±0.7
46±1
TR-EI (both)
657±86
5.6±1.6
0.6±0.7
195±55
TR-SVO (both)
744±128
26.5±12.7
3.2±7.6
349±56
TR-IA (C)
−1,935
0.2
−771
88
TR-IA (NW)
−1,678
1.0
−1,037
28
SRA-EI (both)
416±86
35.9±3.9
8.5±4.2
195±44
Appendix
Table 6. Commons Harvest, seven agents, eight seeds held out from tuning (4–11). Values are means ± 95% Student- t interval half-widths where shown; bold marks the best per column. In parentheses, the metric that selected the setting (C: collective return; NW: Nash welfare; both: the same setting). Sust.: mean step at which apples are eaten. ∗ Means of SRA-IA and SRR-IA (both) and of SRR-EI , SRR-SVO , and SRR-IA+V (NW); half-widths ≤4.2 ( C ), ≤0.6 ( NW ), and ≤1.7 ( L ). L : lowest agent return. Self-scaling control: SVO with ϕ=0 , which only rescales the agent’s own reward.
Env.
Method (selection)
C
NW
L
Behavior
Escape Room
Plain PPO
−0.25±0.39
4.95±0.08
−0.25±0.39
no reliable escape
TR-EI (C)
17.00±0.00
6.96±0.00
−1.00±0.00
concentrated exits, SD 0.49
SRA-EI
15.75±2.95
6.78±0.42
−1.00±0.00
concentrated exits, SD 0.48
SRR-IA + Standing (NW)
13.08±0.28
7.50±0.03
+1.27±0.12
criterion 8/8, SD 0.13
Clean Up
Plain PPO
248±198
37.0±29.8
5.4±4.8
–
TR-SVO (C)
1,405±316
64.6±37.8
5.3±12.5
cleaning share 0.86±0.16
Appendix
Table 7. Escape Room and Clean Up with five agents, eight held-out seeds. C is collective return, NW Nash welfare ( c=5 in Escape Room, 0 in Clean Up), L the lowest agent return. In parentheses, the metric that selected a setting when the two differ. The last column describes behaviour: the standard deviation of the exit rates in Escape Room, about 0.5 when the same agents always exit and 0 when all exit equally often, and in Clean Up the share of the cleaning done by the two lowest earners, 0.40 for an equal split.
Figure 8. Nash welfare over training for the runs of Figure.3 in the main paper. Columns are games and rows are social operators; every learner is at its collective-return selection. Lines are means over eight held-out seeds and shaded regions are pointwise 95% Student- t intervals, with a five-bin moving average.
Figure 9. Lowest agent return over training for the runs, with the same layout and intervals as Figure 8 . The grey line marks zero. Commons Harvest TR-IA falls below the plotted range, to a final mean of −771 , and is marked at the panel edge.
N
M
PPO
TR
Best self-ref.
SRA-EI
≥90%
2
1
0.00
1.00
1.00 ( SRR-EI )
1.00
7/16
3
2
0.00
1.00
1.00 ( SRR-IA )
0.69
3/16
5
3
− 0.01
1.00
1.00 ( SRR-IA )
0.93
3/16
7
4
0.00
1.00
1.00 ( SRR-IA )
0.82
2/16
10
5
0.00
1.00
1.00 ( SRR-IA )
0.69
2/16
12
6
0.00
1.00
1.00 ( SRR-IA )
0.00
2/16
Appendix
Table 8. Escape Room: collective return as a fraction of the optimum, held-out seeds. “Best self-ref.” is the best of 16 self-referenced variants (four operators × four placements) selected on tuning seeds; the last column counts how many of the 16 reach 90% of the optimum.
Figure 10. The value look-ahead. Collective return over training for the reward placement without the look-ahead, SRR , and with it, SRR+V ; plain PPO and the true-reward learner are drawn faintly for reference. Every learner is at its collective-return selection; lines are means over eight held-out seeds and shaded regions, drawn for every learner except plain PPO, are pointwise 95% Student- t intervals; where all seeds agree, as for the true-reward learners at the Escape Room optimum, the interval has no width. Learners below the plotted range are marked at the panel edge with their final means.
Figure 11. Where the self-referenced estimate enters, by group size. Escape Room collective return divided by the optimum for each operator, with the estimate in the policy update, SRA , in the reward, SRR , and in the reward with the value look-ahead, SRR+V , against the true-reward learner. Every cell is tuned separately at its N on tuning seeds; points are means over eight held-out seeds and shaded regions, drawn for SRA and SRR+V , are 95% Student- t intervals.
Figure 12. Final collective return, Nash welfare, and lowest agent return against the number of agents N for plain PPO, TR-EI , SRA-EI , TR-IA , SRR-IA+V , TR-SVO , and SRA-SVO , each at its collective-return selection for that N ; means over eight held-out seeds, bars are 95% t -intervals. The Escape Room optimum grows with N (dotted line). Downward triangles mark means below the plotted range: in Commons Harvest, TR-IA at N=7 ends at −1,935 , and at N=9 TR-IA , SRR-IA+V , SRA-EI , TR-EI , and TR-SVO end at −12,818 , −5,735 , −3,770 , −2,268 , and −1,720 . Commons Harvest Nash welfare and lowest agent return here are the per-step logged values, which differ slightly from the values computed from per-agent means in the other tables.
Figure 13. Who receives the return and who does the work for all nine Clean Up learners with five agents. One block per operator; columns are the true-reward learner, SRA , and SRR+V , each at its collective-return selection. Top row of each block: every agent’s return over training, agents ranked within each seed by their final return and averaged by rank over the eight held-out seeds, darkest the richest. Bottom row: each agent’s final share of the apples, solid, and of the cleaning, dashed, with 95% Student- t intervals; the dotted line is the equal share of 20%.
Method
Setting
C
NW
L
Returns
Plain PPO
–
373±206
104.9±61.8
48.3±45.2
172 / 153 / 48
SRA-EI (NW)
α=2
768±69
245.6±24.5
172.8±27.9
335 / 260 / 173
SRA-SVO
α=7
807±36
258.9±13.5
196.1±23.3
369 / 242 / 196
SRA-IA
0/1
689±32
228.7±10.7
208.6±13.8
257 / 223 / 209
SRR-IA+V
0/0.03
664±213
211.2±74.1
160.3±75.1
296 / 207 / 160
TR-IA
0/0.03
820±37
265.4±16.1
215.7±31.9
361 / 243 / 216
Appendix
Table 9. Clean Up, three agents, eight held-out seeds (95% CIs). IA settings are α/β ; SVO uses ϕ=π/3 . The last column lists the mean return of the highest through the lowest earner.
Method
Setting
C
NW
L
Returns, high → low
Plain PPO
–
313±193
33.7±24.0
5.2±8.2
70 / 60 / 51 / 49 / 47 / 31 / 5
SRA-EI (NW)
α=3
1,953±174
198.5±18.2
34.2±13.7
426 / 409 / 383 / 352 / 285 / 64 / 34
SRA-SVO (NW)
α=5
1,356±111
161.8±13.9
39.0±9.8
266 / 252 / 244 / 239 / 231 / 85 / 39
SRR-IA (NW)
0.003/0.1
1,069±116
151.1±16.2
116.3±17.8
178 / 166 / 162 / 155 / 150 / 142 / 116
SRR-IA+V (NW)
0.003/0.1
1,006±133
140.9±19.8
100.3±26.9
181 / 166 / 154 / 143 / 133 / 129 / 100
TR-IA (NW)
0/0.03
345±296
40.9±37.5
11.4±14.8
76 / 65 / 62 / 55 / 52 / 23 / 11
Appendix
Table 10. Clean Up, seven agents, eight held-out seeds (7–14; 95% CIs). Each learner at the setting selected on the three tuning seeds, by the metric in parentheses. The last column lists the mean return of the highest through the lowest earner.
Figure 14. Who receives the return in Escape Room with five agents, as in Figure 13 . Nash welfare uses the offset c=5 of Eq. ( 15 ).
Figure 15. Who receives the return in Commons Harvest with seven agents, as in Figure 13 . TR-IA collapses and lies below the plotted range.
Figure 16. Agent outcomes under the two tuning objectives. Columns are the three games; rows are the configurations selected on tuning seeds by collective return (top) and Nash welfare (bottom), without the self-scaling controls. Each large panel shows the return of every agent over training: agents are ordered within each seed by their final-window return and then averaged by rank across the eight held-out seeds (darkest: highest final-return rank). The narrow panel shows all eight final seed values and the rank mean ± 95% Student- t interval. Insets give the final C and NW . Ranked curves describe how the return is divided; they do not follow one agent across seeds.
Figure 17. Standing in the repeated game, Escape Room. (a, b) SRR-IA with five agents, without and with standing, the only difference between the two: the fraction of episodes in which each agent exits, ranked within each seed by its final exit rate and averaged by rank over the eight held-out seeds; the dotted line is the even share, two exits in five. (c) Collective return divided by the optimum for the same two runs. (d–g) Lowest agent return over training without and with standing for the four exact pairs. Means over seeds 11–18 with pointwise 95% Student- t intervals. Figure 18 shows every seed.
Figure 18. Every held-out seed of the standing pairs in Escape Room over training: the lowest agent return, top, and the collective return over the optimum, bottom, without standing in grey and with it in colour. Thin lines are single seeds, seeds 11–18, and bold lines their mean. Without standing, every seed settles at a loss of one point for its worst-off agent, and most groups reach the optimum; with standing, the worst-off agent stays above zero on almost every seed, and the group settles at 70 to 80% of the optimum. † The two TR-IA pairs also differ in their PPO settings.
ER( N , M ), with-standing configuration
C/C⋆
Ls
Exit SD
Shared seeds
(3,2) SRA-EI , κ=4
0.711±0.035
1.514±0.220
0.030±0.015
0/8→8/8
(5,3) SRA-EI , κ=4
0.797±0.002
0.114±0.119
0.129±0.003
0/8→6/8
(5,3) SRR-IA , κ=1
0.769±0.016
1.268±0.122
0.131±0.016
0/8→8/8
(5,3) TR-IA , κ=1 †
0.735±0.032
0.942±0.403
0.107±0.020
0/8→8/8
(7,4) SRR-IA , κ=2
0.718±0.039
1.201±0.282
0.104±0.040
0/8→8/8
(7,4) TR-IA , κ=2 †
0.691±0.134
0.476±0.783
0.122±0.050
0/8→6/8
Appendix
Table 11. The four exact pairs of Figure 17 and the two descriptive TR-IA pairs, held-out seeds 11–18. Values for the with-standing arm are means ± 95% Student- t interval half-widths. A seed counts as shared when Ls>0 and the exit rates have a standard deviation below 0.20; the last column gives the shared seeds without → with standing. Every without-standing arm is 0/8. † The TR-IA pairs also differ in their PPO settings.
Figure 19. How many agents pull the lever at the first step of an episode in Escape Room with standing, over 400 evaluation episodes per run, for the four configurations with standing. Bars are means over training seeds and dots single seeds. The line is the binomial distribution that independent choices with the observed pull rate would give; the episode is solved at once only when exactly M agents pull.
Control
Gate
N=3
N=5
N=7
Critic trained on raw returns
off
8.00→7.44
15.91→13.62
21.60→19.97
on
5.98→4.67
13.52→13.03
–
Standing as a policy input
off
8.00→8.00
15.91→15.99
21.60→21.43
on
5.98→6.40
13.52→13.53
–
One shared true ledger
on
5.98→5.70
13.52→13.49
–
with standing input
on
6.40→6.17
13.53→13.54
–
Appendix
Table 12. Collective return of the selected SRA-EI learner in Escape Room, without and with each control, and without or with the standing gate at κ=4 ; three seeds per cell. The optimum is 8, 17, and 26 at N=3 , 5, and 7.
Figure 20. Commons Harvest with recurrent agents and the partial view, trained for 14 million frames; the dotted line marks 7 million, the length of the runs in Figure.4 in the main paper. Rows are operators and columns the collective return, Nash welfare, and lowest agent return; each line is one candidate setting, solid and dashed for the two settings of a learner, means over the eight held-out seeds with pointwise 95% Student- t intervals. One TR-SVO setting had not finished when the figure was made.
Figure 21. Who receives the return in Clean Up when agents 2 and 4 earn two points per apple. Rows are operators and columns the true-reward learner and ours, each at the setting selected under shared rewards. Each line is one agent’s return over training, ranked within each seed from the highest to the lowest earner and averaged by rank over the eight held-out seeds. Our learners reach a higher Nash welfare and leave the poorest agent more than the true-reward learner of the same operator; under SRR-IA+V the agents eat almost equally many apples, so the two agents with the higher value end with about twice the return of the others.
Figure 22. What each learner gives the agents that value apples more, in Clean Up with five agents when agents 2 and 4 value an apple at 2. One column per operator, comparing the true-reward learner with ours, each at the setting selected under shared rewards. Top: return per agent over training for the two high-value agents, solid, and the three others, dashed, means over the eight held-out seeds with pointwise 95% Student- t intervals. Middle and bottom: final apples eaten and share of the cleaning per agent, dark for the high-value agents and light for the others, with 20% an equal share of the cleaning; bars are means with 95% intervals and dots single seeds.
Method
Rewards
Obs./act.
Policies
Critic
Inequity aversion ( Hughes et al., 2018 )
∙
Social value orientation ( McKee et al., 2020 )
∙
Prosocial reward ( Peysakhovich and Lerer, 2018 )
∙
VDN, QMIX ( Sunehag et al., 2018 ; Rashid et al., 2018 )
∙ (team)
∙
∙
COMA ( Foerster et al., 2018b )
∙ (team)
∙
∙
MADDPG ( Lowe et al., 2017 )
∙
∘
∙
Appendix
Table 13. What each method uses from the other agents during training ( ∙ : used; ∘ : in some variants): rewards, observations and actions, policies or gradients, and a centralized critic or joint value. At execution every method listed acts on the agent’s own observation.
EfiArazi School of Computer Science, Reichman University, Herzliya, Israel · Faculty of Computer Science, The College of Management Academic Studies, Rishon LeZion, Israel · School of Communication, Reichman University, Herzliya, Israel