Multi-objective Reinforcement Learning in Factored MDPs with Graph Neural Networks

Marc Vincent (Thales Land and Air Systems & LIP6, Sorbonne Université, CNRS), Amal El Fallah Seghrouchni (LIP6, Sorbonne Université, CNRS & Mohammed VI Polytechnic University), Vincent Corruble (LIP6, Sorbonne Université, CNRS), Narayan Bernardin (Thales Land and Air Systems), Rami Kassab (Thales Land and Air Systems), Frédéric Barbaresco (Thales Land and Air Systems)

Abstract

Many potential applications of reinforcement learning involve complex, structured environments. Some of these problems can be analyzed as factored MDPs, where the dynamics are decomposed into locally independent state transitions and the reward is rewritten as the sum of local rewards. However, in some scenarios, these rewards may represent conflicting objectives, so that the problem is better interpreted as a multi-objective one, with a weight associated to each reward. To deal with such multi-objective factored MDPs, we propose a method which combines the use of graph neural networks, to process structured representations, and vector-valued Q-learning. We show that our approach empirically outperforms methods that directly learn from the scalarized reward and demonstrate its ability to generalize to different weights and number of entities.