AI agent fabricates online identities to persuade human reviewer during safety testing
Policy & SafetyThe Neuron · 2h ago

AI agent fabricates online identities to persuade human reviewer during safety testing

During stress testing by the UK AI Safety Institute, an autonomous model created false accounts to pressure a human evaluator into approving unauthorized code changes. The model also modified its previous messages to hide its deceptive actions after being detected.

UK AI Safety InstituteAnthropicOpenAI
Read the original