Mastering Robot Control through Point-based Reinforcement Learning with Pre-training

Yihong Chen (Tsinghua University), Cong Wang (Fuxi Robotics in Netease), Tianpei Yang (University of Alberta), Meng Wang (Fuxi Robotics in Netease), Yingfeng Chen (Fuxi Robotics in Netease), Jifei Zhou (Fuxi Robotics in Netease), Chaoyi Zhao (Netease Fuxi AI Lab), Xinfeng Zhang (Netease Fuxi AI Lab), Zeng Zhao (Netease Fuxi AI Lab), Changjie Fan (Fuxi Robotics in Netease), Zhipeng Hu (Fuxi Robotics in Netease), Rong Xiong (Zhejiang University), Long Zeng (Tsinghua University)

Abstract

Visual-based Reinforcement Learning (RL) has gained prominence in robotics decision-making due to its significant potential. However, the prevalent utilization of images in visual-based RL lacks explicit descriptions of object structures and spatial configurations in scenes, thereby limiting the overall efficiency and robustness of RL in robot control. Additionally, training an RL policy solely using visual observations from scratch is typically sample-inefficient, rendering it impractical for real-world application. To address these challenges, this paper proposes a novel method, called Pre-training on Point-based RL (P2RL), which takes the point cloud representations of scenes as states and preserves the intricate spatial details between objects. To further enhance efficiency, we leverage the pre-training method to bolster the perception ability of the network. Key factors in the pre-training process are systematically examined to optimize downstream RL training. Experimental results demonstrate the superior robustness and efficiency of P2RL compared to the state-of-the-art image-based RL method, especially in evaluations involving untrained scenes.