Learning Pre-Trained Tacit Behavior for Efficient Multi-Agent Adversarial Coordination
Abstract
In addressing the multi-agent adversarial coordination problem, existing multi-agent reinforcement learning algorithms primarily rely on team-based rewards to guide agent policy updates, often neglecting the utilization of inter-agent relationships, which limits their performance. Drawing inspiration from human tactics, we introduce the concept of tacit behavior to improve the efficiency of multi-agent reinforcement learning by refining the learning process. This paper presents a novel two-phase framework for learning Pre-trained Tacit Behavior for efficient multi-agent adversarial Coordination (PTBC), comprising a tacit pre-training phase and a centralized adversarial training phase. We demonstrate the superiority of our method through comparisons with several algorithms, each of which possesses distinct strengths.