PAC Continuous State Online Multitask Reinforcement Learning with Identification

Yao Liu (Peking University), Zhaohan Guo (Carnegie Mellon University), Emma Brunskill (Carnegie Mellon University)

Abstract

One key feature of a general intelligent autonomous agent is to be able to learn from past experience to improve future performance. In this paper we consider how an agent can leverage prior experience from performing reinforcement learning in order to learn faster in future tasks. We introduce the first, to our knowledge, probably approximately correct (PAC) RL algorithm COMRLI for sequential multitask learning across a series of continuous-state, discreteaction RL tasks. We assume tasks are sampled from a finite number of clusters of Markov decision processes, and provide a bound on the number of steps on which the algorithm makes a suboptimal decision that is substantially smaller on later tasks. We also provide preliminary evidence to suggest our approach may be useful in practice, by showing encouraging simulation performance in a standard domain where it compares favorably to a state-of-the-art algorithm.