Offline Goal-Conditioned Reinforcement Learning with Elastic-Subgoal Diffused Policy Learning

Yaocheng Zhang (Institute of Automation, Chinese Academy of Sciences & School of Artificial Intelligence, University of Chinese Academy of Sciences), Yuanheng Zhu (Institute of Automation, Chinese Academy of Sciences & School of Artificial Intelligence, University of Chinese Academy of Sciences), Yuqian Fu (Institute of Automation, Chinese Academy of Sciences & School of Artificial Intelligence, University of Chinese Academy of Sciences), Songjun Tu (Institute of Automation, Chinese Academy of Sciences & Pengcheng Laboratory), Dongbin Zhao (Institute of Automation, Chinese Academy of Sciences & School of Artificial Intelligence, University of Chinese Academy of Sciences)

Abstract

Goal-conditioned reinforcement learning (GCRL) aims to learn a policy that generalizes across different goal conditions. Compared to non-hierarchical methods, hierarchical GCRL based on subgoals can alleviate the problem of inaccurately estimating the value function for faraway goals in offline learning scenarios, thereby leading to more effective policy learning. Due to the state complexity of the decision-making process, at different states, we require subgoals from varying future time steps to minimize policy errors caused by noisy value functions, rather than using a fixed future time step for selecting subgoals. Therefore, we propose a hierarchical reinforcement learning algorithm with an elastic subgoal steps, called ESD (Elastic Subgoal Diffused Policy Learning). Our method defines a novel high-level policy in which all reachable states surrounding the current state are considered as potential subgoals, and the optimal subgoal is selected among them. Moreover, we use diffusion models to represent the hierarchical policies, enhancing their ability to capture the multimodal data distribution introduced by the elastic subgoal steps and offline data. We evaluate the performance of ESD on multiple goal-conditioned benchmarks, and it demonstrates superior performance compared to previous baselines. Our method effectively reduces the impact of inaccurate value function estimates on policy accuracy, especially in complex tasks and high-dimensional image observations. Code is available at https://github.com/zhyaoch/ESD.