Enhanced Learning from Multiple Demonstrations with a Flexible Two-level Structure Approach
Abstract
Learning from demonstration (LfD) has been emerged as a successful transfer learning technique to speed up reinforcement learning (RL). However, the effectiveness of the LfD heavily depends on the quality of the demonstrations. This work investigates how to enable efficient human-agent (or agent-agent) knowledge transfer and allow the RL agent to extract useful information from multiple demonstrations of different quality. In particular, we aim to avoid the effect of noise or bad examples from the collected demonstration data. Inspired by the multi-armed contextual bandit problem and Human Agent Transfer algorithm, we developed a Flexible Two-level Structured Approach to address the above challenges. Evaluated with Mario, Cart Pole and RC Car domains, the experimental results show that this approach holds the promising capacity to successfully leveraging the demonstrations of different quality.