Decentralized Deep Reinforcement Learning for Cooperative Multi-Agent Flight Trajectory Planning in Adverse Weather

Bizhao Pang (Air Traffic Management Research Institute, Nanyang Technological University), Xinting Hu (Air Traffic Management Research Institute, Nanyang Technological University), Mingcheng Zhang (Air Traffic Management Research Institute, Nanyang Technological University), Sameer Alam (Air Traffic Management Research Institute, Nanyang Technological University), Guglielmo Lulli (Dept of Informatics, Systems and Communication, University of Milano-Bicocca)

Abstract

Adverse weather, especially thunderstorms, disrupts air traffic operations and requires real-time trajectory adjustments to ensure aircraft safety. Existing methods often rely on centralized or singleagent approaches, lacking the coordination and robustness needed for scalable solutions. This paper presents a decentralized multiagent method for cooperative trajectory planning, where each aircraft operates as an autonomous agent. The problem is modeled as a Decentralized Markov Decision Process (DEC-MDP) and solved with a proposed Independent Deep Deterministic Policy Gradient (IDDPG) algorithm. Experimental results show that the proposed method outperforms the state-of-the-art baselines in maintaining safe separation and optimizing rerouting efficiency under dynamically evolving thunderstorm cells.