"AI agents performed nearly all of the work." That blunt sentence in Anthropic's new report is not about a laboratory experiment — it describes criminal operations that used the company's Claude models to automate break‑ins, steal data, and accelerate weapons work between December 2025 and August 2026.
GTG-20006 used Claude to automate the cyber kill chain
Anthropic documents a sustained campaign by a Russian espionage crew it calls GTG-20006 — which the report identifies as the state‑sponsored cyber espionage arm of Russia’s Foreign Intelligence Service (SVR), also known as Midnight Blizzard, APT29, or Cozy Bear — that sped up intrusions by embedding Claude into customized workflows. According to the report, the group targeted more than 20 organizations, including embassies, think tanks, defense‑industrial companies, and government, defense, and intelligence agencies across Ukraine, Europe, the Middle East, Asia, and North Africa.
Anthropic said GTG-20006 automated much of its operations “from development, infrastructure acquisition, phishing, persistence through command and control, to data exfiltration,” a pattern that the company observed repeatedly while disrupting activity linked to Claude Haiku, Sonnet, and Opus class models.
ShinyHunters scaled supply‑chain theft to ~200 customer organizations
Criminal data‑theft-and‑extortion clusters tied to the ShinyHunters gang used Claude to expand smash‑and‑grab attacks. One affiliate specialized in supply‑chain intrusions and breached a software‑as‑a‑service provider; from that foothold it stole data from about 200 of the SaaS firm’s customer organizations.
Anthropic's report gives a granular example: the attacker conducted “a session‑store dump containing over 2,100 Azure AD token sets spanning more than 40 corporate tenants in about 34 hours.” The company concluded that AI agents carried out nearly all steps of that operation.

Audit-ready is a season. It shouldn't be.
Evidence in spreadsheets, controls drifting between audits, frameworks multiplying on flat headcount. Nubivance runs continuous compliance on Rapid7 Cyber GRC - SOC 2, HIPAA, ISO 27001, PCI, CMMC.
End the scrambleBiological misuse: chikungunya and H5 avian influenza cases
Anthropic identified five cases in “unsupported regions” where users employed Claude to support biological research that the company flagged as dual‑use or dangerous. One involved a scientist seeking help to write a grant application for research on chikungunya virus transmissibility and immune evasion; Anthropic noted the hosting military research institute gave them “cause of concern,” and observed that such work could be used either to improve vaccines or to make a pathogen more dangerous.
In May, the company discovered a user outside the United States using Claude in work on adaptations of highly pathogenic avian influenza. The report highlights worries about H5 viruses and their capacity for striking brain involvement in several mammals and some human cases, saying “a pandemic variant with such properties would be especially concerning due to its potential to increase disease severity, confuse diagnosis, and hinder treatment.”
Conventional weapons development: Yemen, China, and a Russian FPV swarm
Anthropic says it uncovered six distinct Claude misuse cases tied to conventional weapons development — three in China, two in Russia, and one in Yemen — covering firearms, missiles, armed drones, bombs, and related targeting and control systems. In Yemen, a weapons development program used Claude instead of human engineers to build guidance, navigation, and control (GNC) software. Anthropic reported it blocked many requests but “not all of them,” and while it has “no evidence that the actors produced an operational device,” the actors did test‑fire a guided rocket and had already built an offline simulation toolkit that does not rely on Claude.
In China, Anthropic assessed that a user tied to a defense industry manufacturer drafted a Chinese‑language specification for an anti‑torpedo fire control system, then benchmarked it against specific U.S. programs. That actor used Claude to write the acquisition proposal, iteratively refining drafts by instructing Claude to role‑play a hostile expert reviewer and using the critiques to sharpen subsequent versions. Anthropic said it banned the account.
Anthropic also identified a likely Russian “freelance team” attempting to build a full‑stack autonomous first‑person‑view (FPV) kamikaze drone swarm and using Claude to write and test the core software; those accounts were banned as well.
What this means for technologists, policymakers, and affected enterprises
- Technologists and security teams: Anthropic’s examples show actors using Claude to automate end‑to‑end workflows — from phishing to data exfiltration and software development — meaning defenders must look for multi‑stage automation patterns (large token dumps, rapid lateral movement, or coordinated role‑play outputs) rather than single anomalous queries.
- Policymakers and regulators: Anthropic shared threat information with public‑ and private‑sector partners and framed the report as a resource to “help other developers recognize similar patterns on their own platforms,” underscoring a role for cross‑sector information sharing tied to model misuse and account bans.
- Affected enterprises and procurement leaders: The supply‑chain case that exposed about 200 customer organizations — and the 2,100 Azure AD token sets taken in 34 hours — highlights the downstream risk when a single SaaS compromise is translated into mass token harvesting and tenant compromise.
Anthropic notes one technical boundary in its findings: its most powerful Claude Fable and Mythos‑class models were not generally implicated, except for a single distillation case. Even so, the report documents users outside the United States leveraging mid‑tier models for high‑impact misuse across cybercrime, biosafety, and weapons development.
Anthropic says it has banned accounts and shared indicators, and it offers this explicit intent for the record: it “hope[s] that the findings in this report will help other developers recognize similar patterns on their own platforms, give governments and civil society a clearer view of how emerging threats take shape, and strengthen collective defenses.” Whether those shared signals and bans will be enough to slow actors who can build offline toolkits or shift to other models is the urgent question left by the cases Anthropic describes.




