Influence Based Reward Shaping in Multiagent Systems
Abstract
Learning based approaches work well to coordinate multiagent systems for a broad range of applications. A key challenge in multiagent learning is that the system reward captures the performance of many agents, making it difficult to determine which agents' actions were helpful. Reward shaping helps address this challenge by isolating the direct impact of an agent's actions. However, when an agent's impact is indirect-such as influencing other teammatesthen existing approaches struggle. Influence based reward shaping addresses indirect impacts by rewarding an agent based on not just its own actions, but also the actions of agents it influenced. Preliminary results demonstrate that this approach leads to better coordination in a guidance mission where leaders must learn to guide followers to points of interest.