Multiagent Q-learning with Sub-Team Coordination

Wenhan Huang (Shanghai Jiao Tong University), Kai Li (Shanghai Jiao Tong University), Kun Shao (Huawei Noah's Ark Lab), Tianze Zhou (Beijing Institute of Technology), Jun Luo (Huawei Noah's Ark Lab), Dongge Wang (EPFL), Hangyu Mao (Huawei Noah's Ark Lab), Jianye Hao (Huawei Noah's Ark Lab), Jun Wang (University College London), Xiaotie Deng (Peking University)

Abstract

For cooperative mutliagent reinforcement learning tasks, we propose a novel value factorization framework in the popular centralized training with decentralized execution paradigm, called multiagent Q-learning with sub-team coordination (QSCAN). This framework could flexibly exploit local coordination within sub-teams for effective factorization while honoring the individual-globalmax (IGM) condition. QSCAN encompasses the full spectrum of sub-team coordination according to sub-team size, ranging from the monotonic value function class to the entire IGM function class, with familiar methods such as QMIX and QPLEX located at the respective extremes of the spectrum. Empirical results show that QSCAN's performance dominates state-of-the-art methods in predator-prey tasks and the Switch challenge in MA-Gym.