Fast Adaptation to External Agents via Meta Imitation Counterfactual Regret Advantage

Mingyue Zhang (Peking University), Zhi Jin (Peking University), Yang Xu (University of Electronic Science and Technology of China), Zehan Shen (Nanjing University), Kun Liu (Peking University), Keyu Pan (University of Electronic Science and Technology of China)

Abstract

This paper focuses on the multi-agent credit assignment problem. We propose a novel multi-agent reinforcement learning algorithm called meta imitation counterfactual regret advantage (MICRA) and a three-phase framework for training, adaptation, and execution of MICRA. The key features are: (1) a counterfactual regret advantage is proposed to optimize the target agents' policy; (2) a meta-imitator is designed to infer the external agents' policies. Results show that MICRA outperforms state-of-the-art algorithms.