The Power of Suggestion
Abstract
Multiagent teams have been shown to be effective in many domains that require coordination among team members. However, finding valuable joint-actions becomes increasingly difficult in tightlycoupled domains where each agent's performance depends on the actions of many other agents. Reward shaping partially addresses this challenge by deriving more "tuned" rewards to provide agents with additional feedback, but this approach still relies on agents randomly discovering suitable joint-actions. In this work, we introduce Counterfactual Agent Suggestions (CAS) as a method for injecting knowledge into an agent's learning process within the confines of existing reward structures. We show that CAS enables agent teams to converge towards desired behaviors more reliably. We also show that improvement in team performance in the presence of suggestions extends to large teams and tightly-coupled domains.