Offline Meta Reinforcement Learning with Weighted Policy Constraints and Proximal Context Collection

Haorui Li (State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Jiaqi Liang (State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS), Linjing Li (State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS & School of Artificial Intelligence, UCAS), Daniel Zeng (State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, CAS & School of Artificial Intelligence, UCAS)

Abstract

Offline meta-reinforcement learning (OMRL) encounters two key challenges: effectively learning the meta-policy from offline datasets and correctly inferring unseen tasks. Existing methods often address the first challenge by imposing policy constraints, but are limited by the suboptimal actions in offline datasets. For the second challenge, most focus on meta-training without enhancing task inference during meta-testing. To address these issues, we propose a novel method called weighted policy conStraints and proximal contExt coLlECtion sTrategy for OMRL (SELECT). During metatraining, we integrate policy constraints with weighted behavior cloning, allowing for more flexible policy learning while maintaining desirable behaviors. In the meta-testing phase, SELECT introduces a proximal context collection strategy that balances exploration and exploitation. This strategy gathers high-quality context, improving task inference and adaptation to unseen tasks. Experimental results show that SELECT significantly reduces the distributional shift, enhances the meta-policy's generalization, and outperforms state-of-the-art methods across various domains.