Provably Efficient Convergence of Primal-Dual Actor-Critic with Nonlinear Function Approximation

Jing Dong (The Chinese University of Hong Kong, Shenzhen), Li Shen (JD Explore Academy), Yinggan Xu (The Chinese University of Hong Kong, Shenzhen), Baoxiang Wang (The Chinese University of Hong Kong, Shenzhen)

Abstract

We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primaldual formulation. Stochastic gradient descent ascent is applied with an adaptive proximal term for robust learning rates. We show the first efficient convergence result with primal-dual actor-critic with a convergence rate of O โˆš๏ธƒ ln(๐‘ ๐‘‘๐บ 2) ๐‘ under Markovian sampling, where ๐บ is the element-wise maximum of the gradient, ๐‘ is the number of iterations, and ๐‘‘ is the dimension of the gradient. Our result is presented with only the Polyak-ลojasiewicz (PL) condition for the dual variable, which is easy to verify and applicable to a wide range of RL scenarios.