Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning

Peter Vamplew (Federation University Australia), Cameron Foale (Federation University Australia), Conor F. Hayes (Lawrence Livermore National Laboratory), Patrick Mannion (University of Galway), Enda Howley (University of Galway), Richard Dazeley (Deakin University), Scott Johnson (Deakin University), Johan Källström (Linköping University), Gabriel Ramos (Universidade do Vale do Rio dos Sinos), Roxana Rădulescu (Vrije Universiteit Brussel / Utrecht University), Willem Röpke (Vrije Universiteit Belgium), Diederik M. Roijers (Vrije Universiteit Brussel)

Abstract

Research in multi-objective reinforcement learning (MORL) has introduced the utility-based paradigm, which makes use of both environmental rewards and a function that defines the utility derived by the user from those rewards. In this paper we extend this paradigm to the context of single-objective reinforcement learning (RL), and outline multiple potential benefits including the ability to perform multi-policy learning across tasks relating to uncertain objectives, risk-aware RL, discounting, and safe RL. We also examine the algorithmic implications of adopting a utility-based approach.