Never Worse, Mostly Better: Stable Policy Improvement in Deep Reinforcement Learning

Pranav Khanna (Indian Institute of Technology), Guy Tennenholtz (Technion), Nadav Merlis (Technion), Shie Mannor (Technion & NVIDIA), Chen Tessler (NVIDIA)

Abstract

The abstract for this paper is not yet available.