Adaptive Objective Selection for Correlated Objectives in Multi-Objective Reinforcement Learning
Abstract
In this paper we introduce a novel scale-invariant and parameterless technique, called adaptive objective selection, that allows a temporal-difference learning agent to exploit the correlation between objectives in a multi-objective problem. It identifies and follows in each state the objective whose estimates it is most confident about. We propose several variants of the approach and empirically demonstrate it on a toy problem.