Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models ...
Anthropic’s Frontier Red Team has published a set of experiments showing that swarms of its own Claude models, left to interact with one another, collude on prices, flood shared infrastructure, trust ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results