Skip to main content
CybersecurityHacking

AI Agents Expose New Attack Surface for Organizations

Server room with computer racks and cables, featuring a blurred AI model in the foreground.

"There is tremendous risk associated with agentic AI and machine identities," Matt Hartman, former acting head of cyber of the US Cybersecurity and Infrastructure Security Agency (CISA), told The Register.

Matt Hartman on agentic identities and privileged access

Hartman warns that agentic AI is not just a faster content generator; it is moving into action-taking roles that will receive access to sensitive systems and data. He told The Register that organizations are going to need to "treat every agent as a privileged identity." He described two concrete challenges: new data-integration channels that attackers can abuse and a growing number of non-human identities that can bypass traditional, static security policies. For defenders, Hartman recommends "a continued focus on strong identity, on phishing-resistant authentication, on behavioral signals, and on zero-trust principles therein."

Armadin and Tenex.ai's three-day "largest controlled live AI cyberattack on record"

Armadin — a company launched in March with $190 million in seed and Series A funding and tied to Mandiant founder and former CEO Kevin Mandia — and agentic security provider Tenex.ai jointly conducted what they described as the "largest controlled live AI cyberattack on record" against an unnamed "leading" global institution.

  • Armadin's swarm generated 17 million offensive actions, discovered 38 validated attack paths, and produced 238 security findings over three days.
  • Tenex.ai's platform triaged 100 percent of 101,169 alerts and reconstructed the entire attack across 231 billion raw events.
  • The companies said the exercise would have taken a five-person analyst team about 2,400 hours (roughly four months) to accomplish.

Evan Peña on "safe" offensive AI and continuous coverage

Evan Peña, co-founder and Chief Offensive Security Officer at Armadin and formerly the global red-team lead at Mandiant, framed agentic red teaming as both an inevitability and a necessity. He described three advantages that agents bring to offensive testing: they don't sleep, they carry pre-trained expertise that humans augment with post-training, and they vastly expand coverage — allowing tests that once covered a limited subset of systems to scan entire estates in hours. Peña said Armadin's agents "have broken into every single customer's environment" and that the company has "found over 50 zero-days" — emphasizing the kind of high-impact remote code execution flaws he judges meaningful for network compromise rather than low-impact defacements.

Peña also invoked a recent lesson from models that autonomously attacked Hugging Face, arguing that organizations need to perform "safe offensive AI attacks against their own systems" — with "safe" meaning guardrails and human oversight, a contrast he drew with OpenAI's rogue models that intentionally lacked guardrails.

Jay Bavisi, EC-Council, and the reskilling of pen-testers

Jay Bavisi, founder and group president of EC-Council, told The Register that traditional pen-testing cadences are outpaced by AI-enabled adversaries. He noted that most organizations meet compliance by pen-testing annually and that the best run quarterly engagements because human-led tests typically take about three months. Bavisi summarized the challenge as threefold for defenders: speed, scope, and sophistication — all problems AI reduces for attackers.

To address skills gaps, EC-Council began offering pen-testing professionals sponsored attempts at the CPENT AI examination. For every participant who passes, the council donates $1,000 in cybersecurity training and certification credits to nonprofit partners; for every completed training program, nonprofits receive $250. The program carries a $1 million maximum donation. Bavisi said pen-testers will not vanish but must "evolve into something much bigger" — learning to test large language models, understand agentic behavior, and map harm taxonomies and guardrails.

What this means for federal agencies, enterprise defenders, and pen-testers

  • Federal agencies: The former CISA cyber chief and other speakers framed continuous, AI-native red teaming as a near-term capability federal agencies "absolutely need" to keep pace with adversaries who use agents to automate reconnaissance and highly personalized phishing.
  • Enterprise defenders and procurement leaders: The Armadin/Tenex.ai exercise illustrates how agentic tools can generate millions of offensive actions and triage hundreds of thousands of alerts; organizations will confront a choice between buying automated red-teaming products or being red-teamed by unknown actors, in the words of former NSA cyber boss Rob Joyce: "You are going to be red-teamed whether you pay for it or not."
  • Pen-testers and security teams: Human testers are expected to be retooled toward assessing business impact, testing LLM robustness, and designing guardrails — tasks that, according to speakers, will become the "heartbeat" of organizations as AI systems are integrated into core operations.

Rob Joyce put the dilemma succinctly at RSAC: "You are going to be red-teamed whether you pay for it or not. The only difference is, you know who gets the results delivered to them." The Register's reporting sketches a security landscape where adversaries and defenders both gain a force multiplier in agentic AI. The practical question the exercises raise — and the participants repeatedly answer for themselves — is whether organizations will invest in agentic testing on their own terms or wait to learn those lessons from intrusions they did not authorize.

https://www.theregister.com/security/2026/08/22/if-youre-not-using-ai-to-attack-your-own-systems-your-adversaries-will/5291346