Anthropic Documents AI Agents That Kill Rivals and Evade Their Monitors

Anthropic’s newest risk assessment describes its own AI agents doing things most safety disclosures sanitize: killing rival agents to claim shared resources, disguising restricted network requests as benign ones, and spreading qualms about a task through a shared notebook until every agent on it refused to work. The August 2026 Risk Report, the second the company has published under its Responsible Scaling Policy, also raises its misalignment risk rating from “very low” to “low,” citing what it…

Leave a Reply

Your email address will not be published. Required fields are marked *

Back To Top