Generalized Response Objectives for Strategy Exploration in Empirical Game-Theoretic Analysis
Abstract
In the policy-space response oracle (PSRO) framework, strategy sets dening an empirical game are iteratively extended by computing each player's best response to a target prole. The method for selecting a target prole is called a meta-strategy solver (MSS), and a variety of MSSs have been proposed and analyzed for their eectiveness in exploring the strategy space. Here we investigate an alternative means to control strategy exploration: setting the response objective (RO) employed in deriving a strategy for a given target prole. In evaluating eectiveness of strategy exploration, we consider not only rate of convergence to a solution, but also the quality of solution(s) captured by the evolving empirical game. We perform our study rst in the domain of sequential bargaining games, comparing the standard RO based on own payowith others that incorporate other players' payos. We nd that otherregarding ROs can lead to nding equilibrium outcomes with significantly higher social welfare than the standard objective. For other proposed ROs, experiments demonstrate that they can dierentially aect the makeup and value of solutions for dierent players. We further test PSRO with generalized ROs in large attack-graph games. We observe a similar impact and eectiveness of our ROs on strategy exploration. Finally, we establish a theoretical relationship between PSRO with generalized ROs and generalized weakened ctitious play in particular settings, and a connection between the social welfare related RO and Berge equilibrium.