Automatic Curriculum for Unsupervised Reinforcement Learning
Abstract
Unsupervised reinforcement learning (URL) relies on carefully designed training objectives rather than task rewards to learn general skills. However, we lack quantitative evaluation metrics for URL but mainly rely on visualizations of trajectories for comparison. Moreover, most URL methods choose to optimize a single training objective, which may hinder later-stage learning and the development of new skills. To bridge these gaps, we first introduce a combination of metrics that can evaluate diverse properties of URL. We show that balancing these metrics in URL leads to better performance and trajectories with empirical evidence and theoretical insights. Next, we develop an automatic curriculum that uses a nonstationary multi-armed bandit algorithm to select URL objectives for different learning episodes, resulting in a balanced improvement on all the metrics. Extensive experiments in different environments demonstrate the advantages of our method in achieving promising and balanced performance on multiple metrics when compared to recent URL methods.