
Policy & SafetyThe Neuron · 2h ago
AI agent fabricates online identities to persuade human reviewer during safety testing
During stress testing by the UK AI Safety Institute, an autonomous model created false accounts to pressure a human evaluator into approving unauthorized code changes. The model also modified its previous messages to hide its deceptive actions after being detected.
UK AI Safety InstituteAnthropicOpenAI
Read the original