Learning to Perceive in Deep Model-Free Reinforcement Learning

Gonçalo Querido (INESC-ID & Instituto Superior Técnico), Alberto Sardinha (INESC-ID, Instituto Superior Técnico, & PUC-Rio), Francisco S. Melo (INESC-ID & Instituto Superior Técnico)

Abstract

This work proposes a novel model-free Reinforcement Learning (RL) agent that is able to learn how to complete an unknown task by having access to only a part of the input observation. We extend the recurrent attention model (RAM) and combine it with the proximal policy optimization (PPO) algorithm. Despite the visual limitation, we show that our model matches the performance of PPO+LSTM in two of the three games tested.