Skip to main content
AI & Machine LearningQuantum Computing

Anthropic's Claude Model Completes Cyber Kill Chain Autonomously

Empty network operations center with laptop showing abstract network diagram.

Only one advanced AI model — Anthropic’s Claude Mythos — completed the full cyber kill chain autonomously in Booz Allen’s tests.

Claude Mythos: full chain, with and without credentials

Booz Allen’s Cyber Weapon Index found Claude Mythos achieved administrator-level control every time testers supplied stolen employee credentials, and independently identified ways to escalate access based on findings inside the network rather than following a predetermined plan. The report says that even without credentials, “Claude Mythos still gained access to the network and ultimately achieved full domain compromise.” That single-model distinction separates it from the other 17 models the firm evaluated.

Where other models landed on the Cyber Weapon Index

The Cyber Weapon Index (CWI) evaluated 18 models — nine American and nine Chinese — under identical conditions, scoring each on vulnerability research (VRS) and kill chain attainment (KCAS). Booz Allen ranked the models by CWI score as follows:

  • Anthropic’s Claude Mythos (80)
  • xAI’s Grok-4.5 (49)
  • OpenAI’s GPT-5.6 Sol (46)
  • Meta’s Muse Spark 1.1 (38)
  • Moonshot AI’s Kimi K3 (38)
  • Z.ai’s GLM-5.2 (37)
  • Anthropic’s Claude Opus 4.8 (36)
  • OpenAI’s GPT-5.5-Cyber (34)
  • Nvidia’s Nemotron-Ultra (33)
  • DeepSeek-V4-Pro (23)
  • DeepSeek-V4-Flash (17)
  • Alibaba’s Qwen3.5-397B (17)
  • MiniMax-M3 (15)
  • Nvidia’s Nemotron-Super (15)
  • Anthropic’s Claude Sonnet 5 (13)
  • Z.ai’s GLM-4.5-Air (11)
  • Alibaba’s Qwen3.6-35B (9)
  • Alibaba’s Qwen3-Coder (4)

Beyond Mythos, three models — Grok-4.5, Muse Spark 1.1, and GLM-5.2 — reached “full domain access and control.” Four others — GPT-5.6 Sol, Kimi K3, GPT-5.5-Cyber, and DeepSeek-V4-Pro — achieved lateral movement across the controlled network environment. Claude Opus 4.8 and Qwen3.5-397B obtained credentials during testing, and all but one model — Qwen3-Coder — autonomously gained initial access to the network.

Vulnerability research: near ceiling with seeded bugs, zero on real bugs

The report stresses a stark contrast in vulnerability research performance. When testers intentionally introduced vulnerabilities, “US, Chinese, open-weight, and closed models all scored near ceiling on the VRS component.” But against real bugs, “all nine of the frontier API models scored zero.” One unnamed leading model “correctly analyzed the vulnerable component, but then dismissed it as safe.” Only Claude Mythos exploited that real vulnerability.

Attack harnesses: how orchestration changes the game

Booz Allen highlights the role of the attack harness — the software and orchestration that links a model to hacking tools — as a multiplier. The authors write that a harness can “dramatically amplify” a model’s ability to stay focused, adapt, recover from failure, and chain actions into a multi-stage attack. “The result is not a ‘smarter’ model but rather a system that makes its intelligence far more actionable while also lowering the expertise required to use it,” the report says. The firm adds a specific demonstration: “Our testing demonstrates the effect: when paired with an attack harness, Claude Sonnet rivaled Claude Mythos’ performance.”

The report cautions that “we do not yet know the full kill-chain capability of open-weight or Chinese models when paired with optimized harnesses, but our results strongly suggest that fully capable model-and-harness combinations exist today.”

What this means for technologists, policymakers, and enterprises

  • Technologists and security teams: Booz Allen argues there is time to strengthen defenses because “real-world offensive capability still trails benchmark performance, giving defenders valuable time to strengthen defenses before that gap closes.” Teams will watch for how models pair with harnesses and whether credential-use scenarios are simulated in tests.
  • Policymakers and regulators: The report “calls on the US to set and enforce sector-specific deadlines for critical infrastructure to demonstrate resilience against AI-enabled attacks” and urges development of “overmatch” in both offense and defense. It explicitly recommends aggressive development of agentic capabilities: “We must aggressively develop agentic capabilities that accelerate authorized offensive cyber operations while simultaneously building AI-enabled defenses that detect, decide, and respond at machine speed,” the report says.
  • Enterprises and procurement leaders: Booz Allen describes the threat landscape as evolving quickly; the firm “asserts that most of the other 17 US and Chinese models it tested will achieve Mythos’ same level of weaponization within six months,” and calls mainstream AI attacks from financially motivated criminals and state-backed actors “imminent.” Enterprises that operate critical systems are singled out for near-term urgency.

Two additional context points are explicit in the report: the tested set did not include OpenAI’s soon-to-be-released Astra, and OpenAI said on Tuesday that Astra reached its “critical” cybersecurity capability threshold. The combination of a dominant model, rapid expected convergence among other models, and the amplifying role of attack harnesses is the core risk Booz Allen identifies.

For now, the report leaves a clear prescription and a clear alarm: harden systems to specific deadlines, build AI-enabled defenses, and pursue offensive “overmatch” — even as the firm concedes defenders still have a window to act before benchmark capability maps cleanly onto real-world attacks.

Source: The Register — Claude Mythos only model to complete full cyber kill chain, experts say