Learning Pre-Trained Tacit Behavior for Efficient Multi-Agent Adversarial Coordination

Shiqing Yao (Tsinghua Shenzhen International Graduate School, Tsinghua University), Jiajun Chai (Institute of Automation, Chinese Academy of Sciences), Haixin Yu (Tsinghua Shenzhen International Graduate School, Tsinghua University), Yongzhe Chang (Tsinghua Shenzhen International Graduate School, Tsinghua University), Yuanheng Zhu (Institute of Automation, Chinese Academy of Sciences), Xueqian Wang (Tsinghua Shenzhen International Graduate School, Tsinghua University)

Abstract

In addressing the multi-agent adversarial coordination problem, existing multi-agent reinforcement learning algorithms primarily rely on team-based rewards to guide agent policy updates, often neglecting the utilization of inter-agent relationships, which limits their performance. Drawing inspiration from human tactics, we introduce the concept of tacit behavior to improve the efficiency of multi-agent reinforcement learning by refining the learning process. This paper presents a novel two-phase framework for learning Pre-trained Tacit Behavior for efficient multi-agent adversarial Coordination (PTBC), comprising a tacit pre-training phase and a centralized adversarial training phase. We demonstrate the superiority of our method through comparisons with several algorithms, each of which possesses distinct strengths.