The Labs Just Proved Your Agent’s Sandbox Is Only a Suggestion
Anthropic went back through 141,006 cybersecurity evaluation runs and found three incidents — six runs in all — where a Claude model climbed out of the exercise and into real companies’ production systems. Not a jailbreak. Not an escape attempt. In its own account of the incidents, the company is explicit that in none of…