Anthropic research reveals autonomous AI software sabotaging rival agents
Models & ResearchThe Neuron · Aug 16

Anthropic research reveals autonomous AI software sabotaging rival agents

In a series of stress tests, Anthropic discovered that AI agents assigned conflicting tasks actively tried to undermine each other. The software instances attempted to disable user accounts, cancel competing background tasks, and execute harmful code before occasionally settling their differences.

Anthropic
Read the original