Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning

Willem Röpke (Vrije Universiteit Brussel), Mathieu Reymond (Mila - Quebec Artificial Intelligence Institute & Vrije Universiteit Brussel), Patrick Mannion (University of Galway), Diederik M. Roijers (City of Amsterdam & Vrije Universiteit Brussel), Ann Nowé (Vrije Universiteit Brussel), Roxana RadulescuXXX (Utrecht University & Vrije Universiteit Brussel)

Abstract

An important challenge in multi-objective reinforcement learning is obtaining a Pareto front of policies to attain optimal performance under di"erent preferences. We introduce Iterated Pareto Referent Optimisation (IPRO), which decomposes# nding the Pareto front into a sequence of constrained single-objective problems. This enables us to guarantee convergence while providing an upper bound on the distance to undiscovered Pareto optimal solutions at each step. We evaluate IPRO using utility-based metrics and its hypervolume and# nd that it matches or outperforms methods that require additional assumptions. By leveraging problem-speci#c single-objective solvers, our approach also holds promise for applications beyond multi-objective reinforcement learning, such as planning and path#nding.