Improving the Effectiveness of Potential-based Reward Shaping in Reinforcement Learning

Henrik Müller (L3S Research Center), Daniel Kudenko (L3S Research Center)

Abstract

Potential-based reward shaping is often used to incorporate prior knowledge of how to solve the task into reinforcement learning because it can formally guarantee policy invariance. In this work, we highlight the dependence of effective potential-based reward shaping on the initial Q-values and external rewards, which determine the agent's ability to exploit the shaping rewards to guide its exploration and achieve increased sample efficiency. We formally derive how a simple linear shift of the potential function can be used to improve the effectiveness of reward shaping without changing the structure of the potential function and thus its implicitly encoded preferences, and without having to adjust the initial Qvalues. We verify our theoretical findings on tabular Q-learning and demonstrate the application of our findings in deep reinforcement learning.