Dual-Policy-Guided Offline Reinforcement Learning with Optimal Stopping

Weibo Jiang (Tsinghua Shenzhen International Graduate School, Tsinghua University), Shaohui Li (Tsinghua Shenzhen International Graduate School, Tsinghua University), Zhi Li (Tsinghua Shenzhen International Graduate School, Tsinghua University), Yuxin Ke (Tsinghua Shenzhen International Graduate School, Tsinghua University), Zhizhuo Jiang (Tsinghua Shenzhen International Graduate School, Tsinghua University), Yaowen Li (Tsinghua Shenzhen International Graduate School, Tsinghua University), Yu Liu (Department of Electronics, Tsinghua University)

Abstract

Policy-guided offline reinforcement learning (POR) decomposes the offline reinforcement learning (offline RL) problem into goal estimation and goal-conditioned execution subproblems, leading to improved performance. However, we reveal that the preciseness of the estimated goal massively affects the performance and robustness of the trained goal-conditioned policy. To overcome this problem, we propose an offline RL model with dual guide-policies to improve the preciseness of the goal and reduce the variance. The proposed dual-policy-guided offline RL (Dual POR) adopts an integrating function, which balances the goals predicted by two guide-policies to obtain a refined goal. Moreover, we employ the optimal stopping strategy to schedule the training process, which dramatically shortens the training process and improves the generalization. The proposed Dual POR achieves state-of-the-art performance on the D4RL datasets with reduced variances. The improvements in highcomplexity tasks are even significant, which indicates the potential of the proposed Dual POR in real-world applications.