Cooperative Multi-Agent Reinforcement Learning in Convention Reliant Environments

Jarrod Shipton (University of the Witwatersrand Johannesburg)

Abstract

Multi-Agent Reinforcement Learning (MARL) has seen a move towards creating algorithms which can be trained to work cooperatively with partners. Typically MARL is done in self-play (SP). Recent works show agents trained with SP often achieve near optimal results when paired with one another, however, they form arbitrary play conventions which can perform poorly when mismatched. This led to research into algorithms which have been developed to form strategies which avoid the need for convention matching and allow for zero-shot coordination (ZSC) with any novel partner. ZSC solves the problem of convention matching, and is useful in short interactions, however in prolonged or repeated interaction this comes at the cost of optimality. Avoiding conventions leaves the challenge of being unable to exploit known, existing conventions and achieve higher levels of optimality. In this work we use population training with a belief of the partner type to exploit conventions which could exist, leading to high rewards over prolonged interactions. We demonstrate that our method is able to better adapt in convention reliant environments over repeated interactions than current state-of-the-art competing ZSC methods.