Experience-replay Innovative Dynamics

Tuo Zhang (University of Birmingham), Leonardo Stella (University of Birmingham), Julian Barreiro-Gomez (Khalifa University)

Abstract

Multi-agent reinforcement learning (MARL) has achieved groundbreaking success in recent years. Yet, several open problems remain, including nonstationarity and instability. Evolutionary game theory (EGT) provides a theoretical framework to tackle instability by leveraging the properties of its most well-known model, namely, the replicator dynamics, for theoretical guarantees of convergence to Nash equilibria. However, these guarantees do not hold true in certain settings, e.g., zero-sum games. In contrast, innovative dynamics, such as the Brown-von Neumann-Nash (BNN) or Smith, retain the convergence guarantees in these settings. We develop a novel MARL algorithm based on innovative dynamics with a sampling process that resembles experience replay. We show that our approach is theoretically grounded as other state-of-the-art MARL algorithms, but most importantly it outperforms other approaches in the case of nonstationary environments.