Mitigating Non-Stationarity in Deep Reinforcement Learning with Clustering Orthogonal Weight Modification

Guoqing Ma (Institute of automation, Chinese Academy of Sciences & School of Future Technology, University of Chinese Academy of Sciences), Yuhan Zhang (Institute of automation, Chinese Academy of Sciences & School of Future Technology, University of Chinese Academy of Sciences), Yuming Dai (Institute of automation, Chinese Academy of Sciences & School of Future Technology, University of Chinese Academy of Sciences), Guangfu Hao (Institute of automation, Chinese Academy of Sciences & School of Future Technology, University of Chinese Academy of Sciences), Yang Chen (Institute of Automation, Chinese Academy of Sciences & Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Chinese Academy of Sciences), Shan Yu (Institute of Automation, Chinese Academy of Sciences, School of Future Technology, University of Chinese Academy of Sciences, and Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Chinese Academy of Sciences)

Abstract

RL agents often operate under the assumption of environmental stationarity, which poses a great challenge to learning efficiency since many environments are inherently non-stationary in state distribution. To address this issue, we introduce the Clustering Orthogonal Weight Modified (COWM) layer, which can be integrated into the policy network of any RL algorithm and mitigate non-stationarity effectively. By employing clustering techniques and a projection matrix, the COWM layer stabilize the learning process. Empirically, the COWM layer is integrated into various RL methods and outperforms state-of-the-art methods on the DMControl benchmark, highlighting its robustness and generality across various tasks and algorithms.