Dual-Policy-Guided Offline Reinforcement Learning with Optimal Stopping
Abstract
Policy-guided offline reinforcement learning (POR) decomposes the offline reinforcement learning (offline RL) problem into goal estimation and goal-conditioned execution subproblems, leading to improved performance. However, we reveal that the preciseness of the estimated goal massively affects the performance and robustness of the trained goal-conditioned policy. To overcome this problem, we propose an offline RL model with dual guide-policies to improve the preciseness of the goal and reduce the variance. The proposed dual-policy-guided offline RL (Dual POR) adopts an integrating function, which balances the goals predicted by two guide-policies to obtain a refined goal. Moreover, we employ the optimal stopping strategy to schedule the training process, which dramatically shortens the training process and improves the generalization. The proposed Dual POR achieves state-of-the-art performance on the D4RL datasets with reduced variances. The improvements in highcomplexity tasks are even significant, which indicates the potential of the proposed Dual POR in real-world applications.