The shift follows internal testing that suggests human intervention is often less reliable than automated safeguards. In a study of 1,053 paid users, auto mode successfully identified 89% of harmful actions. By comparison, human reviewers caught only 13.6% of potential issues, a discrepancy Anthropic attributes to users habitually approving 97% of permission prompts without careful examination.
Boris Cherny, head of Claude Code, stated that his team has used auto mode exclusively for months, noting that returning to manual permission prompts is no longer a viable workflow. Under the new standard, the system proceeds automatically unless it identifies an action as destructive, irreversible, or attempting to access data outside the user's defined environment. To mitigate risks, the company is deploying additional safety measures, including prompt injection screening and customizable hard deny rules to prevent unauthorized data exfiltration.

Comments (0)
No comments yet. Be the first!