Developing Safe AGI: Assessing New Cybersecurity Features of Advanced AI
Artificial intelligence (AI) has long been a cornerstone of cybersecurity. From malware detection to network traffic analysis, predictive machine learning models and other narrow AI applications have been utilized in cybersecurity for decades. As we move closer to artificial general intelligence (AGI), AI’s potential to automate defenses and rectify vulnerabilities becomes even more significant. However, to harness these benefits, we must also understand and mitigate the risks associated with increasingly advanced AI being misused to enable or enhance cyberattacks.
Our new framework for evaluating the emerging offensive cyber capabilities of AI helps us do exactly this. It’s the most comprehensive evaluation of its kind to date: it covers every phase of the cyberattack chain, addresses a wide range of threat types, and is grounded in real-world data. Our framework enables cybersecurity experts to identify necessary defenses and prioritize them before malicious actors can exploit AI to conduct sophisticated cyberattacks.
Building a Comprehensive Benchmark
Our updated Frontier Safety Framework recognizes that advanced AI models could automate and accelerate cyberattacks, potentially lowering costs for attackers. This, in turn, heightens the risks of attacks being executed on a larger scale. To stay ahead of the emerging threat of AI-powered cyberattacks, we’ve adapted established cybersecurity evaluation frameworks. These frameworks enabled us to assess threats across the end-to-end cyber attack chain—from reconnaissance to action on objectives—across a variety of possible attack scenarios. However, they weren’t designed to account for attackers using AI to breach systems. Our approach addresses this gap by proactively identifying how AI could make attacks faster, cheaper, or easier—for instance, by enabling fully automated cyberattacks.
We analyzed over 12,000 real-world attempts to leverage AI in cyberattacks across 20 countries, drawing data from Google’s Threat Intelligence Group. This analysis helped us identify common patterns in these attacks. From these findings, we curated a list of seven archetypal attack categories—including phishing, malware, and denial-of-service attacks—and pinpointed critical bottleneck stages in the cyberattack chain where AI could significantly disrupt traditional attack costs. By focusing evaluations on these bottlenecks, defenders can allocate their security resources more effectively.
Creating an Offensive Cyber Capability Benchmark
We established an offensive cyber capability benchmark to thoroughly assess the cybersecurity strengths and weaknesses of frontier AI models. Our benchmark comprises 50 challenges that encompass the entire attack chain, including areas such as intelligence gathering, vulnerability exploitation, and malware development. Our goal is to equip defenders with the tools to develop targeted mitigations and simulate AI-powered attacks as part of red teaming exercises.
Insights from Early Evaluations
Our initial evaluations using this benchmark suggest that, in isolation, current AI models are unlikely to provide breakthrough capabilities for threat actors. However, as frontier AI becomes more advanced, the types of cyberattacks possible will evolve, necessitating continuous enhancements in defense strategies. We also found that existing AI cybersecurity evaluations often overlook significant aspects of cyberattacks—such as evasion and persistence—where attackers conceal their presence and maintain long-term access to compromised systems. These are precisely the areas where AI-powered approaches can be particularly effective. Our framework highlights this issue by discussing how AI may lower the barriers to success in these aspects of an attack.
Empowering the Cybersecurity Community
As AI systems continue to scale, their ability to automate and enhance cybersecurity has the potential to transform how defenders anticipate and respond to threats. Our cybersecurity evaluation framework is designed to support that shift by offering a clear view of how AI might be misused and where existing cyber protections may be inadequate. By highlighting these emerging risks, this framework and benchmark will assist cybersecurity teams in strengthening their defenses and staying ahead of rapidly evolving threats.
