Safe Systems with Unsafe Agents: Challenges and Opportunities
Abstract
Generative AI (genAI) models have advanced significantly in recent years, enabling artificial cognitive agents to process information and interact with their environments. While much effort has focused on aligning genAI models to produce reliable behavior, less attention has been given to their safe integration into critical systems. This work draws parallels between human safety practices and genAI agent safety, proposing a shift from individual agent alignment to a system-level perspective. We identify key weaknesses in genAI powered agents, connect these to established human safety errors, and explore how these vulnerabilities manifest in critical systems. Building on existing research in system safety, we outline mitigation strategies that encompass not only model-level improvements but also cognitive structures and system interfaces, opening a new avenue of research into cognitive genAI agent safety.