Learning in Non-Stationary MDPs as Transfer Learning

M. M. Hassan Mahmud (University of Edinburgh, UK)

Abstract

In this paper we introduce the MDP-with-agents model for addressing a particular sub-class of non-stationary environments where the learner is required to interact with other agents. The behaviorpolicies of the agents are determined by a latent variable that changes rarely, but can modify the agent policies drastically when it does change (like traffic conditions in a driving problem). This unpredictable change in the latent variable results in non-stationarity. We frame this problem as transfer learning in a particular subclass of MDPs, which we call MDPs-with-agents (MDP-A), where each task/MDP requires the learner to learn to interact with opponent agents with fixed policies. Across the tasks, the state and action space remains the same (and is known) but the agent-policies change. We transfer information from previous tasks to quickly infer the combined agent behavior policy in a new task after some limited initial exploration, and hence rapidly learn an optimal/nearoptimal policy. We propose a transfer algorithm which given a collection of source behavior policies, eliminates, using a novel statistical test, the policies that do not apply in the new task in time polynomial in the relevant parameters. We also perform experiments in three interesting domains and show that our algorithm significantly outperforms relevant alternative algorithms.