Search-Improved Game-Theoretic Multiagent Reinforcement Learning in General and Negotiation Games

Zun Li (DeepMind & University of Michigan), Marc Lanctot (DeepMind), Kevin R. McKee (DeepMind), Luke Marris (DeepMind), Ian Gemp (DeepMind), Daniel Hennes (DeepMind), Kate Larson (University of Waterloo & DeepMind), Yoram Bachrach (DeepMind), Michael P. Wellman (University of Michigan), Paul Muller (DeepMind)

Abstract

Multiagent reinforcement learning (MARL) has benefited significantly from population-based and game-theoretic training regimes. One approach, Policy-Space Response Oracles (PSRO), employs standard reinforcement learning to compute response policies via approximate best responses and combines them via meta-strategy selection. We augment PSRO by adding a novel search procedure with generative sampling of world states, and introduce two new meta-strategy solvers based on the Nash bargaining solution. We evaluate PSRO's ability to compute approximate Nash equilibrium, and its performance in negotiation games: Colored Trails and Dealor-no-Deal. We conduct behavioral studies where human participants negotiate with our agents (𝑁 = 346). Search with generative modeling finds stronger policies during both training time and test time, enables online Bayesian co-player prediction, and can produce agents that achieve comparable social welfare negotiating with humans as humans trading among themselves.