Exploiting Causal Structure for Transportability in Online, Multi-Agent Environments
Abstract
Autonomous agents may encounter the transportability problem when they suffer performance deficits from training in an environment that differs in key respects from that in which they are deployed. Although a causal treatment of transportability has been studied in the data sciences, the present work expands its utility into online, multi-agent, reinforcement learning systems in which agents are capable of both experimenting within their own environments and observing the choices of agents in separate, potentially different ones. In order to accelerate learning, agents in these Multi-agent Transport (MAT) problems face the unique challenge of determining which agents are acting in similar environments, and if so, how to incorporate these observations into their policy. We propose and compare several agent policies that exploit local similarities between environments using causal selection diagrams, demonstrating that optimal policies are learned more quickly than in baseline agents that do not. Simulation results support the efficacy of these new agents in a novel variant of the Multi-Armed Bandit problem with MAT environments.