Reward-Machine-Guided, Self-Paced Reinforcement Learning
Abstract
Self-paced reinforcement learning (RL) aims to improve the sample efficiency of RL by automatically creating sequences, i.e., curricula, of probability distributions over contexts. However, existing selfpaced RL methods fail in tasks that involve temporally extended behaviors. As a remedy, we exploit prior knowledge about the underlying task structure and develop a self-paced RL algorithm guided by reward machines, i.e., a finite-state machine that encodes such structure. The proposed algorithm integrates reward machines in the updates of 1) the policy and value functions obtained by an RL algorithm, and 2) the automated curriculum that generates context distributions. Our empirical results evidence that the proposed algorithm achieves optimal behavior in cases where existing methods fail, and also reduces curriculum length and variance.