Interleaved Q-Learning with Partially Coupled Training Process
Abstract
This paper studies estimating the maximum expected value (MEV) of several independent random variables (RVs). No unbiased estimator exists without knowing the distributions of those RVs a priori. Two of the most famous estimators, maximum estimator (ME) and double estimator (DE), yield positive bias and negative bias respectively. We propose a coupled estimator (CE) which subsumes ME and DE as special cases and yields a bias between that of ME and DE, while maintaining the same variance bound. Furthermore, a simple yet effective variance reduction technique is proposed and verified in the experiments. The instantiated algorithm in the Markov decision process (MDP) setting, called interleaved Q-learning, outperforms Q-learning and double Q-learning in some highly stochastic environments. Insights on how to adapt the coupling ratio in CE and hence make interleaved Q-learning automatically shift between Q-learning and double Q-learning are provided and verified in the experimental section.