Status-quo Policy Gradient in Multi-Agent Reinforcement Learning

Pinkesh Badjatiya (Microsoft), Mausoom Sarkar (Adobe), Nikaash Puri (Adobe), Jayakumar Subramanian (Adobe), Abhishek Sinha (Waymo), Siddharth Singh (University of Maryland), Balaji Krishnamurthy (Adobe)

Abstract

Individual rationality, which involves maximizing expected individual returns, does not always lead to high-utility individual or group outcomes in multi-agent problems. For instance, in multi-agent social dilemmas, Reinforcement Learning (RL) agents trained to maximize individual rewards converge to a low-utility mutually harmful equilibrium. In contrast, humans evolve useful strategies in such social dilemmas. Inspired by ideas from human psychology that attribute this behavior to the status-quo bias, we present a status-quo loss (π‘†π‘„πΏπ‘œπ‘ π‘ ) and the corresponding policy gradient algorithm that incorporates this bias in an RL agent. We demonstrate that agents trained with π‘†π‘„πΏπ‘œπ‘ π‘  learn high-utility policies in several social dilemma matrix games (Prisoner's Dilemma, Matching Pennies, Chicken Game). To apply SQLoss to visual input games where cooperation and defection are determined by a sequence of lower-level actions, we present GameDistill, an algorithm that reduces a visual input game to a matrix game. We empirically show how agents trained with SQLoss on GameDistill reduced versions of Coin Game and Stag Hunt learn high-utility policies. Finally, we show that π‘†π‘„πΏπ‘œπ‘ π‘  extends to a 4-agent setting by demonstrating the emergence of cooperative behavior in the popular Braess' paradox.