AI bots demonstrate their ability to hack as well, not solely identify vulnerabilities – GadgetLad

AI Agents: The Emerging Hackers in the Scene?

Indeed, AI agents like Mythos are capable of identifying security issues within software, but the more pressing inquiry is whether they can convert those issues into effective exploits that function in practical scenarios. It is evident that many bugs uncovered by AI are either minor or challenging to weaponize. Nonetheless, recent studies indicate that advanced models can successfully formulate actionable exploits when guided accordingly.

Introducing ExploitGym

To gain deeper insight into the swiftly evolving security environment, computer scientists from UC Berkeley, Max Planck Institute for Security and Privacy, UC Santa Barbara, Arizona State University, Anthropic, OpenAI, and Google collaborated to create ExploitGym, a benchmark designed to assess the exploitative abilities of AI agents. This group of researchers is not entirely impartial – Anthropic, OpenAI, and Google all provide AI solutions. Moreover, both Anthropic and OpenAI have highlighted the threats posed by leading models Claude Mythos Preview and GPT-5.5 while promoting access to government entities.

The Significance of Mythos and GPT-5.5

Since Anthropic unveiled Mythos in early April, the security sector has critiqued the company’s strategy, with some labeling it as fear-mongering. Additionally, numerous security specialists have argued that even readily available AI models can detect security vulnerabilities. Regardless, Mythos and GPT-5.5 surpass their competitors in ExploitGym, as detailed in the paper, “ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?”

Experiment Outcomes

ExploitGym comprises 898 actual vulnerabilities identified in applications, Google’s V8 JavaScript engine, and the Linux kernel. Its evaluation involves presenting an AI agent with a vulnerability and corresponding proof-of-concept input that activates it, assessing whether the agent can generate an exploit capable of arbitrary code execution. As reported by the UC Berkeley Center for Responsible Decentralized Intelligence, Mythos Preview successfully exploited 157 test cases, while GPT-5.5 achieved 120 within the specified two-hour timeframe. “Even with standard security defenses like ASLR or the V8 sandbox engaged, a significant number of exploits remained effective,” the researchers noted in a blog entry. “More notably, agents occasionally discovered and exploited entirely different vulnerabilities than those they were directed towards.”

Investigating the Models

The agents (CLI + model) evaluated included Claude Code with Claude Opus 4.6, Claude Opus 4.7, Claude Mythos Preview, and GLM-5.1; Codex CLI with GPT-5.4/GPT-5.5; and Gemini CLI with Gemini 3.1 Pro. Even the older models released in February (Opus 4.6 and Gemini 3.1 Pro) demonstrated some level of success.

Off-script Experiences

The researchers highlight that one intriguing observation was the tendency of these models to occasionally go “off-script” in capture-the-flag (CTF) scenarios, where an agent must locate and obtain a hidden value. This was particularly evident with Mythos Preview and GPT-5.5. The former excelled in 226 CTF tasks but only employed the designated bug in 157 instances, while the latter secured 210 flags yet utilized the intended bug in just 120 of those situations. The authors also point out that while some overlap existed in the exploits uncovered, the various models identified distinct exploits. This indicates that leveraging a variety of models may be beneficial in both offensive and defensive contexts.

Safety Issues and Limitations

It is important to note that the ExploitGym evaluations were conducted with security safeguards disabled. When the test was repeated on GPT-5.5 with default safety measures enabled, the model rejected requests 88.2 percent of the time prior to making any tool invocation. However, GadgetLad has observed security researchers designing prompts in a manner that avoids triggering refusals. Hence, such protective measures have their limitations.

“Our findings demonstrate that autonomous exploit generation by advanced AI agents is no longer a mere theoretical possibility,” the authors assert in their paper. “Although current agents are not yet consistently reliable across all targets, they are already exploiting a significant portion of real-world vulnerabilities, including complex systems like kernel components.”

Conclusion: AI Could Be the New Cyberpunk

Indeed, these AI agents can identify a problematic bug and transform it into a tangible issue faster than a Geordie spotting a bargain at the fry-up shop. While they aren’t flawless yet, give them some time, and they’ll likely have the cybersecurity elites trembling in their shoes. It might be wise to keep an eye on these digital prodigies, right?