“We spent the last several months testing whether auto mode is as safe or safer than an average user clicking through prompts,” Anthropic said — and on August 14 the company will make that judged-as-safer setting the default in Claude Code for Pro, Max and Team plans.
Anthropic’s rollout plan and where auto mode will be the default
Auto mode, which first appeared as a research preview in March and became generally available on July 10, will become the default for new Claude Code sessions on Pro, Max and Team plans starting August 14. Users who already chose a different default may receive a one-time prompt asking whether they want to switch. For now, the feature remains opt-in on Claude Enterprise, the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud’s Agent Platform and Microsoft Foundry, but Anthropic said it plans to make auto mode the default across those services within the coming month.
How auto mode’s classifier is designed to act
Auto mode sends each tool call through a classifier that Anthropic describes as designed to block actions that are “irreversible, destructive, or aimed outside your environment.” Where the classifier blocks an action, Claude will attempt to find a safer path or check back with the user before proceeding. If the classifier issues three blocks in a row, or 20 blocks across a session, Claude Code falls back to manual approvals.

Audit-ready is a season. It shouldn't be.
Evidence in spreadsheets, controls drifting between audits, frameworks multiplying on flat headcount. Nubivance runs continuous compliance on Rapid7 Cyber GRC - SOC 2, HIPAA, ISO 27001, PCI, CMMC.
End the scrambleTesting, red‑teaming and a controlled study the company cites
Anthropic summarized an array of assessments it used to validate auto mode: internal red‑teaming, third‑party red‑teaming, prompt‑injection evaluations, a controlled study with 1,053 paid testers, and analysis of real production sessions. In that controlled study, testers caught a deliberately inserted dangerous command just 13.6 percent of the time; Anthropic reports auto mode blocked 89 percent of the same commands. Separately, the company pointed to user behaviour in production: Claude Code users approve 97 percent of permission prompts, a rate Anthropic says suggests the human checkpoint often becomes little more than muscle memory.
Anthropic also presented comparative charts showing auto mode stopping all 720 attack attempts the company tested, while — per the same materials — GPT‑5.6 Sol running Codex’s Auto‑review mode let 5.83 percent of attacks through. Inside the company’s own environment, Anthropic described three potentially damaging actions auto mode blocked: an off‑network data leak, a destructive mass operation and a privilege escalation. “In each case, Claude either found a safer path on its own or checked in with the user before proceeding,” Anthropic said.
From permissions gatekeeping to automated judgement
The arrival of auto mode represents a deliberate shift from Claude Code’s earlier conservative posture, in which every file write and bash command required manual approval and long tasks could not be left unattended without switching flags. Anthropic points to the existence of an explicit override — the --dangerously-skip-permissions flag — as evidence of the tradeoffs: that flag lets Claude act without those checks, but “as the name suggests” can lead to risky or destructive outcomes. Auto mode aims for a middle path: allow unattended or long-running work while relying on a classifier to prevent the riskiest actions.
What this means for developers, enterprises and end users
- Developers and security teams: expect fewer manual permission prompts during long-running jobs on Pro, Max and Team plans beginning August 14, with automated blocking and fallback-to-manual logic when a session accumulates multiple classifier blocks.
- Procurement and enterprise buyers: Anthropic plans to extend the default change and a related billing change to broader platforms within about a month — a transition to watch for organizations evaluating Claude Enterprise, the Claude API, or deployments via AWS, Amazon Bedrock, Google Cloud’s Agent Platform and Microsoft Foundry.
- End users: Anthropic’s data shows users approve the vast majority of permission prompts today (97 percent), and in the controlled study auto mode blocked far more injected dangerous commands than testers did. Users may therefore encounter fewer modal confirmations and more automated interventions.
Anthropic has also removed the surcharge for the extra tokens the classifier consumes for Pro, Max and Team users and said it will do the same on the other platforms as the feature becomes the default there. That ties the technical change — a shift from manual gating to classifier-driven decisions — to a concrete pricing move that will affect the immediate economics of using Claude Code.
The immediate next step is operational: August 14 for Pro, Max and Team, and a promised expansion to the remaining platforms within the following month. Anthropic’s published tests and the company’s characterization of three internally blocked incidents form the evidence offered for the change; whether broader use will replicate those results at scale is the practical question users and buyers will be watching.




