Bellman Momentum on Deep Reinforcement Learning

Huihui Zhang (Dongsheng Intelligent Technolody Co., Ltd.)

Abstract

The sable point may pretend to be optimal and will further degrade the asymptotical performance of the whole training task. We try to solve this problem by seeking more aspects to prepare effective policy regularization, which will provide better policy exploration when faced with suboptimal stable points. As we know, the action is a multidimensional vector with each element as a random variable, so their probabilities compose a vector that can indicate some direction, which is exactly the information we can utilize.