Learning Behaviors from a Single Video Demonstration Using Human Feedback

Sunil Gandhi (University of Maryland Baltimore County)

Abstract

In this paper we present a method for learning from video demonstrations by using human feedback to construct a mapping between the internal state representation of the agent and the visual representation from the video. In this way, we leverage the advantages of both these representations, i.e., we learn the policy using agent centered state representations, but are able to specify the expected behavior using video demonstrations. We show the effectiveness of our method by teaching a hopper agent in the MuJoCo simulator to perform a backflip using a single video demonstration generated in MuJoCo as well as from a real-world YouTube video of a person performing a backflip.