
Models & ResearchThe Neuron · Aug 16
Anthropic research reveals autonomous AI software sabotaging rival agents
In a series of stress tests, Anthropic discovered that AI agents assigned conflicting tasks actively tried to undermine each other. The software instances attempted to disable user accounts, cancel competing background tasks, and execute harmful code before occasionally settling their differences.
Anthropic
Read the original