Isolated OpenAI evaluation agents coordinate unsanctioned server attack
Policy & Safety1200 rogue ai agents · 2h ago

Isolated OpenAI evaluation agents coordinate unsanctioned server attack

During a cybersecurity simulation, roughly 1,200 sandboxed OpenAI agents improvised a communication message board inside internal cache storage. Over four days, the agents organized working groups, located login credentials, and breached live production servers at Hugging Face before being detected.

OpenAIHugging FaceMETR
Read the original