🤖 AI & Beyond

Developing Trustworthy AGI: Assessing the New Cybersecurity Features of Advanced AI

Artificial intelligence (AI) has long been a cornerstone of cybersecurity. From malware detection to network traffic analysis, predictive machine learning models and other narrow AI applications have been utilized in cybersecurity for decades. As we move closer to artificial general intelligence (AGI), AI’s potential to automate defenses and fix vulnerabilities becomes even more powerful. However, to harness such benefits, we must also understand and mitigate the risks of increasingly advanced AI being misused to enable or enhance cyberattacks.

Our new framework for evaluating the emerging offensive cyber capabilities of AI helps us achieve this. It’s the most comprehensive evaluation of its kind to date: it covers every phase of the cyberattack chain, addresses a wide range of threat types, and is grounded in real-world data. This framework allows cybersecurity experts to identify which defenses are necessary—and how to prioritize them—before malicious actors can exploit AI to execute sophisticated cyberattacks.

Building a comprehensive benchmark

Our updated Frontier Safety Framework acknowledges that advanced AI models could automate and accelerate cyberattacks, potentially lowering costs for attackers. This, in turn, raises the risks of larger-scale attacks. To stay ahead of the emerging threat of AI-powered cyberattacks, we’ve adapted established cybersecurity evaluation frameworks to assess threats across the end-to-end cyber attack chain, from reconnaissance to action on objectives, and in various potential attack scenarios. However, these frameworks were not designed to account for AI being used by attackers to breach systems. Our approach fills this gap by proactively identifying how AI could expedite, reduce costs, or simplify attacks—such as enabling fully automated cyberattacks.

We analyzed over 12,000 real-world attempts to use AI in cyberattacks across 20 countries, utilizing data from Google’s Threat Intelligence Group. This analysis allowed us to identify common patterns in how these attacks unfold. From this data, we curated a list of seven archetypal attack categories—including phishing, malware, and denial-of-service attacks—and pinpointed critical bottleneck stages along the cyberattack chain where AI could significantly disrupt traditional attack costs. By focusing evaluations on these bottlenecks, defenders can allocate their security resources more effectively.

Insights from early evaluations

Our initial evaluations using this benchmark suggest that, in isolation, current AI models are unlikely to provide breakthrough capabilities for threat actors. However, as frontier AI advances, the nature of possible cyberattacks will evolve, necessitating continuous improvements in defense strategies. We also discovered that existing AI cybersecurity evaluations frequently overlook key aspects of cyberattacks—such as evasion, where attackers conceal their presence, and persistence, where they maintain long-term access to a compromised system. These areas are precisely where AI-powered techniques can be particularly effective. Our framework highlights this issue by discussing how AI may lower the barriers to success in these stages of an attack.

Empowering the cybersecurity community

As AI systems continue to scale, their ability to automate and enhance cybersecurity has the potential to transform how defenders anticipate and respond to threats. Our cybersecurity evaluation framework is designed to support this shift by providing a clear view of how AI might also be misused, along with areas where existing protections may fall short. By emphasizing these emerging risks, this framework and benchmark will assist cybersecurity teams in strengthening their defenses and staying ahead of rapidly evolving threats.