Mini-batch Bayesian Inverse Reinforcement Learning for Multiple Dynamics

Abstract

Inverse reinforcement learning is a method that estates a reward function from experts demonstrations. Most existing inverse reinforcement learning methods assume that an expert gives demonstrations in a fixed environment, although the expert can provide demonstrations for a specific objective in multiple environments. In such cases, normal practice is to use demonstrations in multiple environments to estimate the expert's reward. Herein, we formulate this problem based on a Bayesian inverse reinforcement learning framework and propose a mini-batch Markov chain Monte Carlo method. An advantage of our method is scalability. Our proposed method is scalable with respect to a number of environments in which expert demonstrations are generated. Experimental results show quantitatively that the proposed method outperforms existing inverse reinforcement learning methods.