Exploration in the Face of Parametric and Intrinsic Uncertainties
Abstract
In distributional reinforcement learning (RL), the estimated distribution of the value functions model both the parametric and intrinsic uncertainties. We propose a novel, efficient exploration method for Deep RL that has two components. The first is a decaying schedule to suppress the intrinsic uncertainty. The second is an exploration bonus calculated from the upper quantiles of the learned distribution. In Atari 2600 games, our method achieves 483 % average gain in cumulative rewards over QR-DQN.