Tech

Anthropic shifts Claude Code to auto mode default for commercial plans

Starting August 14, 2026, Pro, Max, and Team subscribers will run autonomous coding sessions by default, with Anthropic citing data that human reviewers approve 97% of permission prompts.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
AI safety classifier blocks dangerous commands as reflexive user approvals rise

Anthropic has announced that Claude Code will operate in auto mode by default for Pro, Max, and Team subscription plans beginning August 14, 2026. The shift is designed to facilitate longer-running autonomous development work while implementing a safety classifier to intercept hazardous commands. This change addresses internal data indicating that developers frequently approve permission prompts reflexively, with a 97% approval rate for individual requests.

The new default configuration routes each tool call through a classifier aimed at blocking actions that are irreversible, destructive, or directed outside the user’s environment. When the system identifies a potential risk, it typically attempts to find a safer alternative or requests explicit user confirmation. If the classifier blocks an action three times consecutively or twenty times within a session, the system falls back to manual approvals. For the affected plans, Anthropic is waiving the token overhead previously charged for running the safety classifier.

Anthropic’s decision follows extensive testing, including red-teaming and a controlled study involving 1,053 paid testers. The research found that auto mode blocked 89% of dangerous commands, compared to a 13.6% detection rate by human participants. Furthermore, the classifier’s block rate remained consistent regardless of session length, whereas human vigilance dropped significantly after 50 or more prior prompts. The company also collaborated with Apollo Research to harden the classifier against synthetic adversarial attacks, reducing the miss rate on a held-out test set from 12% to 7%.

While the change applies immediately to Pro, Max, and Team users, auto mode remains opt-in for Enterprise, API, and cloud platform users, including those on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry. This allows administrators time to review the update and configure managed settings. Anthropic plans to extend the default setting to all remaining channels in the coming month, alongside a further waiver of classifier costs for those segments.

Early adopters such as Adobe, Nuro, Gusto, and Garner Health have already integrated auto mode into their production workflows. Internal metrics suggest that teams using auto mode ship approximately 25% more pull requests than those relying on manual review. Anthropic maintains that while the system reduces risk, it does not eliminate it, and still recommends human review for high-stakes changes to production infrastructure.

Continue reading

More from Tech

Read next: Open-source tool claims 97 per cent token savings for AI agents
Read next: Valvoline Unveils August 2026 Promotional Offers for Service and Retail Buyers
Read next: Developer Antirez releases native MiniMax H3 inference engine for Apple Silicon