Improving Generalization with Cross-State Behavior Matching in Deep Reinforcement Learning

Guan-Ting Liu (National Taiwan University), Guan-Yu Lin (National Taiwan University), Pu-Jen Cheng (National Taiwan University)

Abstract

Representation learning on visualized input is an essential yet challenging task for deep reinforcement learning (RL). To help the RL agent learn more general and discriminative representation among various states, we present cross-state self-constraint (CSSC). This novel technique regularizes representation learning by comparing state embedding similarities across different state-action pairs. We test our proposed method on the OpenAI Procgen benchmark with Rainbow and PPO and demonstrate significant improvement across most Procgen environments.