The Evolutionary Dynamics of Soft-Max Policy Gradient in Multi-Agent Settings

Martino Bernasconi (Politecnico di Milano), Federico Cacciamani (Politecnico di Milano), Simone Fioravanti (Gran Sasso Science Institute), Nicola Gatti (Politecnico di Milano), Francesco Trovò (Politecnico di Milano)

Abstract

We investigate the mean dynamics of the soft-max policy gradient algorithm in multi-agent normal-form games by resorting to evolutionary game theory and dynamical system tools. First, we consider the best-response problem analysis, where a single learner plays against a fixed opponent in continuous time. For such dynamics, we provide a complete characterization of the set of bad initializations (points for which the dynamics initially move towards sub-optimal strategies). Then, we resort to models based on single-and multipopulation games, showing that the dynamics preserve the volume and, in arbitrary instances, it is impossible to obtain last-iterate convergence when the equilibrium of the game is fully mixed. Furthermore, we give empirical evidence that dynamics starting from close initial points may expand over time, thus showing that the behavior of the dynamics in games with fully-mixed equilibrium is chaotic.