Gaze Supervision for Mitigating Causal Confusion in Driving Agents
Abstract
Imitation Learning (IL) algorithms show promise in learning humanlevel driving behavior, but they often suffer from "causal confusion, " a phenomenon where the lack of explicit inference of the underlying causal structure can result in misattribution of the relative importance of scene elements, especially pronounced in complex scenarios like urban driving with abundant information per time step. Our key idea is that while driving, human drivers naturally exhibit an easily obtained, continuous signal that is highly correlated with causal elements of the state space: eye gaze. We collect human driver demonstrations in a CARLA-based VR driving simulator, allowing us to capture eye gaze in the same simulation environment commonly used in prior work. Further, we propose a method to use gaze-based supervision to mitigate causal confusion in driving IL agents-exploiting the relative importance of gazed-at and notgazed-at scene elements for driving decision-making. We present quantitative results demonstrating the promise of gaze-based supervision improving the driving performance of IL agents.