β-DQN: Improving Deep Q-Learning By Evolving the Behavior

Hongming Zhang (University of Alberta & Amii), Fengshuo Bai (Shanghai Jiao Tong University), Chenjun Xiao (CUHK-Shenzhen), Chao Gao (Edmonton Research Center, Huawei), Bo Xu (CASIA), Martin Müller (University of Alberta & Amii)

Abstract

While many sophisticated exploration methods have been proposed, their lack of generality and high computational cost often lead researchers to favor simpler methods like 𝜖-greedy. Motivated by this, we introduce 𝛽-DQN, a simple and efficient exploration method that augments the standard DQN with a behavior function 𝛽. This function estimates the probability that each action has been taken at each state. By leveraging 𝛽, we generate a population of diverse policies that balance exploration between state-action coverage and overestimation bias correction. An adaptive meta-controller is designed to select an effective policy for each episode, enabling flexible and explainable exploration. 𝛽-DQN is straightforward to implement and adds minimal computational overhead to the standard DQN. Experiments on both simple and challenging exploration domains show that 𝛽-DQN outperforms existing baseline methods across a wide range of tasks, providing an effective solution for improving exploration in deep reinforcement learning.