Agent-Time Attention for Sparse Rewards Multi-Agent Reinforcement Learning
Abstract
Sparse and delayed rewards pose a challenge to single agent reinforcement learning. This challenge is amplified in multi-agent reinforcement learning (MARL) where credit assignment of these rewards needs to happen not only across time, but also across agents. We propose Agent-Time Attention, a neural network model with auxiliary losses for redistributing sparse and delayed rewards in collaborative MARL. We provide a simple example to demonstrate how providing agents with their own local redistributed rewards over shared global redistributed rewards leads to better policies.