Joint Intrinsic Motivation for Coordinated Exploration in Multi-Agent Deep Reinforcement Learning
Abstract
Multi-agent deep reinforcement learning (MADRL) often struggles to learn strongly coordinated tasks, as performance depends not only on one agent's behavior but rather on the joint behavior of multiple agents. In this context, a group of agents can benefit from actively exploring different joint strategies to determine the most efficient one. In this paper, we propose an approach for rewarding strategies where agents collectively exhibit novel behaviors. We present JIM (Joint Intrinsic Motivation), a multi-agent intrinsic motivation method that rewards joint trajectories based on a centralized measure of novelty. We show how JIM can be used to improve state-of-the-art MADRL methods in a highly coordinated task, demonstrating the crucial role of coordinated exploration.