SCRIMP: Scalable Communication for Reinforcement- and Imitation-Learning-Based Multi-Agent Pathfinding

Yutong Wang (National University of Singapore), Bairan Xiang (National University of Singapore), Shinan Huang (National University of Singapore), Guillaume Sartoretti (National University of Singapore)

Abstract

In this paper, we propose SCRIMP, a multi-agent reinforcement learning approach for multi-agent path finding. Our method learns individual policies from very small FOVs (3x3), by relying on a highly-scalable global/local communication mechanism based on a modified transformer. We further introduce a state-value-based tiebreaking strategy to improve performance in symmetric situations and intrinsic rewards to encourage exploration while mitigating the long-term credit assignment problem. Empirical evaluations indicate that SCRIMP can outperform other state-of-the-art learningbased planners with larger FOVs and even yield similar performance as a classical centralized planner.