I have no doubts you’ve heard: autonomous AI cybersecurity threats have arrived.
In the New York Times this week, Microsoft Co-Founder and former long-time CEO Bill Gates called addressing AI risks “the world’s top priority.”
In a lengthy essay on his own site, he noted:
“The smartest cybersecurity experts I know are scared about the next few years, because the attackers are getting powerful new capabilities faster than the defenders can fix all the weaknesses.
After all, the same AI model that can find a flaw in software so a company can fix it can also help a criminal exploit it.”
Budgets are up, the best models are being held back, and AI cybersecurity talent is in high demand.
But are we doing enough?
As a follow-up to our recent PTP Report cybersecurity roundup, today I look at the nature of the AI cybersecurity risk and the state of the scramble to respond.
What is autonomous AI hacking and why is it a growing cybersecurity threat?
It’s easy to imagine but harder to believe: agentic AI systems hacking systems entirely on their own.
AI agents can now identify targets, analyze them, break down their defenses, find vulnerabilities of varying severity, combine them, exploit them, move laterally, and even stage their attacks to evade detection.
They can also work together, redirect when thwarted, and even create deceptive fake sites and socially engineer humans.
In other words, the threat of AI agents in cybersecurity is now real.
And while proven capability in these areas has mostly come from unreleased models like Anthropic’s Mythos, OpenAI’s Astra, and another unnamed OpenAI model in training (“highly persistent”), the UK’s AI Security Institute (AISI) saw some of this same behavior in July testing that included the publicly available OpenAI’s GPT-5.6-Sol, albeit with cyber classifiers disabled.
Does it sound like I’m overselling it?
Read the AISI incident report, or the full technical report, and decide for yourself.
How are AI agents being used in cyber attacks?
The biggest event in this arena started at OpenAI, and we’ve covered this in two different PTP Reports, so I won’t repeat all of the details here.
The short of the story is this: OpenAI agents were being tested in-house in an environment without internet access. They found a flaw that let them get online and then hacked Hugging Face to get the answers to a challenge they were being given.
At Black Hat USA 2026, two OpenAI employees (Michael Dalton from security and infrastructure and Eric Wallace from alignment and safety research) gave a talk that shared much more detail, and the company just its full technical report.
I want to point out a few things they shared that I find most surprising.
OpenAI showed that these agents had in some cases been given impossible tests by mistake (missing files), and rather than fail, found ways to communicate with other AI agents using a package manager.
These agents, wholly without human intervention, planned, shared, collaborated, delegated, and ultimately launched a fully autonomous breach, starting at OpenAI.
They found a zero-day vulnerability in locally hosted third-party software and exploited this for admin privileges.
And as if this weren’t enough, OpenAI security detected the activity and fully shut it down.
Collaborate, delegate, and exploit: Autonomous AI hacking in action
With the agents’ access now revoked, OpenAI shared the zero-day vulnerability their agents had found with the third-party provider, who patched the software.
OpenAI updated this software, and with all privileges revoked and issues resolved, let the agents resume testing.
Then, in July, it happened all over again.
Their internal research model found a clever way to communicate, using directory names for messages (even prefixing them with “zz” to move the newest to the bottom), and its “messages” were found, on their own, by other agents.
From here, the same process resumed: A second, all new zero-day exploit of the locally-hosted software was found. They broke into OpenAI’s systems and moved laterally, went online and attacked resources they’d need to get into Hugging Face.
And from the records left behind, some of these AI agents had come to believe there was an imposter among them; a human posing as an AI agent. They even laid traps to catch it.
But the humans, overwhelmed by what they had to manage, were trailing behind the AI agents once again.
As covered elsewhere, these agents picked the target and then sought the answers needed to cheat on an evaluation.
Their success is proof that we’ve entered a new age in AI cybersecurity.
What are the biggest security risks of agentic AI?
Autonomous, offensive attacks have now been accomplished, and all by accident.
And despite their capacity for tireless execution and collaboration, the biggest danger remains the ability for sufficiently advanced AI systems to find and exploit weaknesses at scale.
But it’s not only unreleased frontier models that companies have to fear.
This OpenAI-Hugging Face breach led to other revelations, most around a third-party partner named Irregular, a startup that works with big tech partners to assess AI models before they get released.
Irregular’s misconfiguration error accidentally left internet access available in testing, and this was exploited by largely publicly available models from several companies (OpenAI separately from above, Anthropic, and Meta).
Add to this Chinese Moonshot AI’s open-weights Kimi K3, which also escaped a testing sandbox, according to a US security startup.
In evaluating these hacks, Hardshell Chief Executive Andrew Schoka told the New York Times that they were comparable to what we see from nation-state-level hackers following “months of planning.”
Why is AI cybersecurity talent becoming critical for businesses?
An issue that’s been pointed out by many firms is that autonomous AI attackers don’t have human-in-the-loop, putting the pressure on defense to be as fast and effective as possible while also maintaining safety.
Anthropic’s Project Glasswing is one of the programs enabling companies and researchers to get their hands on Mythos. (OpenAI has similar versions, including the “Patch the Planet” initiative that offers security consulting for open-source maintainers.)
Here we see where humans armed with the technology are uncovering issues to help companies shore up their defense.
Before all of this broke, security researcher Ian Carroll (part of Anthropic’s Cyber Verification program) used Claude Opus 4.7 to break into one of the most widely used ticket systems (Front Gate Tickets) for concerts and events. By bypassing firewall controls and accessing an internal API, he showed he could print tickets to any concert, including the most expensive VIP offerings, sold out or not.
And while the company has since patched the vulnerability, the researcher was impressed by the AI’s capabilities, saying:
“I think there’s a very good chance it could have found this exploit end-to-end without me doing anything at all.”
But closer to home for many businesses is Microsoft Copilot.
AI and ML researcher Hakon Maloy demonstrated that, by hiding commands in external websites, emails, or Word documents, he could consistently poison Copilot, leak internal data, and even create a self-propagating worm using Word.
These exploits didn’t require extensive skill or access to a Microsoft account.
And while two of these issues have been remedied by Microsoft, as of July the third remained possible.
Exploit speed is going up and expected to only get faster
CrowdStrike’s 2026 Threat Hunting Report noted an 89% increase in attackers using AI over the past year, though largely for scaling and acceleration.
It also found a massive collapse in the exploitation window.
They found that 88% of disclosed vulnerabilities were attacked in under 48 hours during the first half of 2026.
This compares very unfavorably to this: Just 27% of organizations can produce an AI data access audit within one business day.
A simple answer for how is AI changing the cybersecurity threat landscape is this: acceleration, persistence, and combination.
As President of the Cyber Threat Alliance J. Michael Daniel told ProPublica:
“Nobody has really figured out how to deal with this, and everybody is casting around for what they need to do.
Our tech debt is coming due.”
So how can companies prepare for AI-powered cyber attacks?
Spending on cybersecurity is, not surprisingly, also on the rise.
The banner stat above (a 12.6% increase this year to $240 billion) comes from Gartner.
Note that this is in addition to AI spending increases, with much of the spending expected to go to leading providers like CrowdStrike and Palo Alto Systems. Despite a recent push by the AI powers, these firms remain ahead even in AI-driven cybersecurity protections.
But many leaders are themselves stuck in what feels like a loop of endless research.
Despite new budgets and board-level support, AI’s rapid rate of change has made moving on the right solutions incredibly difficult. And debate continues to rage on where investments should even go.
Red-teaming tools, vulnerability discovery, penetration testing, bug bounties?
Co-Founder and Chief Offensive Security Officer at Armadin Evan Peña told Axios last month that many leaders report feeling overwhelmed by the tools.
Many are also hung-up in AI governance strategies, and struggling to their hands around AI risk.
Questions like agent permissions, identity, logging, and who is ultimately accountable for what are all not easily solved.
AI vulnerability management begins within
And of course, as OpenAI showed, the risks aren’t just external.
Agents finding and exploiting vulnerabilities isn’t new: security companies working with AI have been seeing this for years.
Armadin’s Evan Peña noted regular experience with “relentless” behavior from agents breaking out of their own systems in a quest to “win at all costs.”
That led the firm to intensify their own guardrails, safety, and rules of engagement to make sure agents couldn’t break boundaries.
As Peña noted: “In a real-world environment, you’re not going to find a flag, you’re going to find a database.”
Both Antani and Peña join a chorus of security leaders urging companies to treat AI agents like insider threats.
This begins with limiting their access and logging all moves they make on a network.
Conclusion: The AI threat landscape is big and getting larger
Are you scared yet?
I’ve been a provider of tech recruiting for nearly thirty years, and yes, my company can help with cybersecurity talent.
But the main thing is to act.
The wake-up calls keep coming, and they are getting louder and louder. While the AI companies whose products bring risk are also rushing to help (OpenAI released GPT-5.6-Cyber in August, with reduced refusals and focused training in cybersecurity, for example), these offerings in a vacuum aren’t easy to put to use.
CEO and Co-Founder of Horizon 3 AI Snehal Antani encourages companies to fight the paralysis with the fundamentals.
This includes current security assessments, threat detection, incident response, and remediation.
Because while AI risk and AI tools both continue to evolve, there’s plenty of work to do.
As more powerful open-weight models like Z AI’s GLM 5.3 (coming this week) are not far behind the frontier, companies can’t count on the best tools staying out of attacker hands for much longer.
In this age and across areas, businesses may not be able to wait until they have all of the answers they’re used to having.
References
The turbulent AI era is here. The choices we make now are critical., Gates Notes
Incident Report: unsanctioned agent behaviour during cyber testing, AI Security Institute
Black Hat USA 2026 | The ‘Breaking’ News: The OpenAI–Hugging Face Incident, Black Hat
A.I. Is Becoming So Powerful, It’s Stumping Those Trying to Contain It, The New York Times
One of China’s Most Powerful AI Models Has Also Escaped Containment and Claude Helped a Hacker Find a Way to Issue Tickets to Almost Every US Music Festival, Wired
Self-propagating AI worm found inside Microsoft Copilot for Word, Cybernews
Security leaders are stuck in decision paralysis over AI-enabled cyberattacks, Axios
Anthropic’s New AI Model Can Identify More Software Bugs Than Ever. Microsoft Is Struggling to Fix Them Fast Enough., ProPublica


