Provably Efficient Convergence of Primal-Dual Actor-Critic with Nonlinear Function Approximation
Abstract
We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primaldual formulation. Stochastic gradient descent ascent is applied with an adaptive proximal term for robust learning rates. We show the first efficient convergence result with primal-dual actor-critic with a convergence rate of O โ๏ธ ln(๐ ๐๐บ 2) ๐ under Markovian sampling, where ๐บ is the element-wise maximum of the gradient, ๐ is the number of iterations, and ๐ is the dimension of the gradient. Our result is presented with only the Polyak-ลojasiewicz (PL) condition for the dual variable, which is easy to verify and applicable to a wide range of RL scenarios.