A Theory of Mind Approach as Test-Time Mitigation Against Emergent Adversarial Communication
Abstract
Multi-Agent Systems (MAS) is the study of multi-agent interactions in a shared environment. Cooperative Multi-Agent Reinforcement Learning (CoMARL) is a learning framework that leverages cooperative mechanisms or policies that exhibit cooperative behavior. Explicitly, there are works on learning to communicate messages from CoMARL agents; however, non-cooperative agents have been shown to learn sabotage a cooperative team's performance through adversarial communication messages. To address this issue, we propose a technique which leverages local formulations of Theory-of-Mind (ToM) to distinguish exhibited cooperative behavior from non-cooperative behavior before accepting messages from any agent. We demonstrate the efficacy and feasibility of the proposed technique in empirical evaluations in a centralized training, decentralized execution (CTDE) CoMARL benchmark.