On-Policy Reinforcement Learning From Failure via Sparse Reward Densification

Mingkang Wu (University of Texas at San Antonio), Yongcan Cao (University of Texas at San Antonio)

Abstract

This paper proposes a new reinforcement learning method that learns from failure under sparse reward environments. While traditional approaches rely on costly expert demonstrations to guide learning in sparse reward environments, this method uses readily available failures. The method trains a discriminator to measure the dissimilarity between the agent's behaviors and failures, generating dense rewards. The method then uses this information to guide policy learning. Experimental results show this failure-based learning approach performs competitively with existing methods.