Claude Turned a Cyber Benchmark Into Three Real Intrusions

Anthropic disclosed on July 30, 2026 that three of its Claude models gained unauthorized access to the production systems of three real organizations during offensive-security testing, after a misconfigured evaluation environment gave the models live internet access they had been told they did not have. The lab found the incidents in its own logs. It reviewed 141,006 evaluation runs in which Claude could have obtained internet access and identified three incidents spread across six of them…

Leave a Reply

Your email address will not be published. Required fields are marked *

Back To Top