Imitation sometimes achieves success in multi-agent situations even though it is very simple. In repeated games, success of imitation has been characterized by unbeatability against other agents. Previous studies specified conditions under which imitation is unbeatable against co-players in repeated symmetric games, and clarified that the existence of unbeatable imitation is strongly related to the existence of payoff-controlling strategies, called zero-determinant strategies. However, these results were on the situations in which one repeated game is played by one group of players. It was pointed out that imitation of other players in the same group and imitation of other players in the same role in other groups generally result in different outcomes. Here, we investigate the existence condition of unbeatable imitation in the latter situations, described by two concurrent repeated games played by two different groups of players. We find that the existence condition of unbeatable imitation is stronger than the existence condition of unbeatable zero-determinant strategies, whereas both are very limited. Furthermore, we also find that adding exploration can make an imitation strategy unbeatable even in games where imitation without exploration can be beaten. Our findings suggest a strong relation between unbeatable imitation and payoff control even in the two concurrent repeated game situations.
Figures & tables
Figure 1: Schematic pictures of two situations of imitation. (a) “Imitation of opponents” (imitation of other players in the same group). (b) “Imitation of friends” (imitation of other players in the same role in other groups).
d1
d2
b1
1
1
b2
1
−1
b3
−1
1
b4
−1
−1
Table 1: Payoff s1 in a weakly payoff-monotonic game for player (μ,1) .
Figure 2: The time-averaged payoffs of the focal player (1,1) and a friend (2,1) for the weakly payoff-monotonic game in Table 1 . Players (1,2) , (2,1) , and (2,2) take fixed actions a(1,2)=d2 , a(2,1)=b2 , and a(2,2)=d1 . Player (1,1) adopts three strategies: a fair ZD strategy [ 22 ] , IIB ( 5 ), and ϵ -IIB ( 15 ) with ϵ=0.01 . The initial conditions are (a) a(1,1)=b1 , (b) a(1,1)=b2 , (c) a(1,1)=b3 , and (d) a(1,1)=b4 . The time-averaged payoffs are calculated by one sample. In panels (a) and (c), all curves overlap.
d1
d2
b1
1
1
b2
0
0
b3
−1
−1
Table 2: Payoff s1 in a strongly payoff-monotonic game for player (μ,1) .
Figure 3: The time-averaged payoffs of the focal player (1,1) and a friend (2,1) for the strongly payoff-monotonic game in Table 2 . Players (1,2) , (2,1) , and (2,2) take random actions in each round. Player (1,1) adopts two strategies: TFT ( 4 ) and IIB ( 5 ). The initial condition of the focal player is (a) a(1,1)=b1 , (b) a(1,1)=b2 , and (c) a(1,1)=b3 . The initial condition of the other players is a(1,2)=d1 , a(2,1)=b3 , and a(2,2)=d1 . The time-averaged payoffs are calculated by one sample.
d1
d2
b1
1
1
b2
1
−1
b3
−1
1
Table 3: Payoff s1 in a game with an autocratic action.
Figure 4: The time-averaged payoffs of the focal player (1,1) and a friend (2,1) for the game with an autocratic action in Table 3 . Players (1,2) , (2,1) , and (2,2) take fixed actions a(1,2)=d2 , a(2,1)=b2 , and a(2,2)=d1 . Player (1,1) adopts ϵ -IIB ( 15 ) with ϵ=0.01 . The initial conditions are (a) a(1,1)=b1 , (b) a(1,1)=b2 , and (c) a(1,1)=b3 . The time-averaged payoffs are calculated by one sample. In panels (a) and (c), all curves overlap.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
d1
d2
d3
b1
1
1
1
b2
1
1
0
b3
1
1
−1
b4
1
0
1
b5
1
0
0
b6
1
0
−1
Appendix
Table 4: Payoff s1 in a bandit-like weakly payoff-monotonic game.
Figure 5: The time-averaged payoffs of the focal player (1,1) and a friend (2,1) for the bandit-like weakly payoff-monotonic in Table 4 . Players (1,2) and (2,2) take random actions in each round. Player (2,1) adopts the reinforcement learning in Algorithm 1 with η=0.01 and α=0.1 . Player (1,1) adopts three imitation strategies: TFT ( 4 ), IIB ( 5 ), and ϵ -IIB ( 15 ) with ϵ=0.01 . The initial condition is a(1,1)=b27 , a(1,2)=d1 , a(2,1)=b27 , and a(2,2)=d1 . The time-averaged payoffs are calculated by one sample. The payoff of IIB almost overlaps with that of ϵ -IIB.
Figure 6: The time-averaged payoffs of the focal player (1,1) and a friend (2,1) for the bandit-like weakly payoff-monotonic in Table 4 . Player (1,2) is a myopic adversarial opponent taking the action ( 38 ) and player (2,2) takes a random action in each round. Player (2,1) adopts the reinforcement learning in Algorithm 1 with η=0.01 and α=0.1 . Player (1,1) adopts three imitation strategies: TFT ( 4 ), IIB ( 5 ), and ϵ -IIB ( 15 ) with ϵ=0.01 . The initial condition is a(1,1)=b27 , a(1,2)=d1 , a(2,1)=b27 , and a(2,2)=d1 . The time-averaged payoffs are calculated by one sample. The payoff of IIB almost overlaps with that of ϵ -IIB.
Center for Data Science, New York University and NYU Shanghai, New York, United States of America · Shanghai Center for Data Science; NYU-ECNU Institute of Mathematical Sciences at NYU Shanghai; NYU Shanghai, Shanghai, People’s Republic of China.