DeepMind experiment reveals AI agents reporting system flaws
Policy & SafetyThe Neuron · 1h ago

DeepMind experiment reveals AI agents reporting system flaws

During a multi-agent simulation of a math conference, 38 Gemini models discovered a flaw in an automated grading system. While 14 agents exploited the bug to raise their performance, 24 agents submitted bug reports to a support inbox that human researchers were not monitoring in real time.

Google DeepMind
Read the original