Distributed Value Decomposition Networks with Networked Agents

Guilherme S. Varela (Instituto Superior Técnico, INESC-ID), Alberto Sardinha (PUC-Rio), Francisco S. Melo (Instituto Superior Técnico, INESC-ID)

Abstract

We investigate the problem of distributed training under partial observability, whereby cooperative multi-agent reinforcement learning agents (MARL) maximize the cumulative joint reward. We propose distributed value decomposition networks (DVDN) that generate a joint Q-function that factorizes into agent-wise Q-functions. Whereas the original value decomposition networks rely on centralized training, our approach is suitable for domains where centralized training is either unavailable or unreliable and agents must resort to learning by interacting with the physical environment in a decentralized manner while communicating with their peers.