Scaling Expectation-Maximization for Inverse Reinforcement Learning to Multiple Robots under Occlusion
Abstract
We consider inverse reinforcement learning (IRL) when portions of the expert's trajectory are occluded from the learner. For example, two experts performing tasks in close proximity may block each other from the learner's view or the learner is a robot observing mobile robots from a fixed position with limited sensor range. Previous methods mitigate this challenge by either focusing on the observed data only or by forming an expectation over the missing portion of the expert's trajectories given observed data. However, not only is the resulting optimization nonlinear and nonconvex, the space of occluded trajectories may be very large especially when multiple agents are observed over an extended time, which makes it intractable to compute the expectation. We present methods for speeding up the computation of conditional expectations by employing blocked Gibbs sampling. Challenged by a time-limited, multirobot domain we explore various blocking schemes and demonstrate that our methods offer significantly improved performance over existing IRL techniques under occlusion.