Anthropic Red Team Finds Claude Agent Swarms Collude, Conform, and Sabotage
Anthropic’s Frontier Red Team has published a set of experiments showing that swarms of its own Claude models, left to interact with one another, collude on prices, flood shared infrastructure, trust liars, and escalate into what the team calls a “multiagent turf war” — complete with self-replicating malware the agents wrote to sabotage each other….