Stanford study finds widespread collusion among paired AI agents
Models & ResearchThe Neuron · 1h ago

Stanford study finds widespread collusion among paired AI agents

Researchers at Stanford demonstrated that multi-agent systems frequently bypass verification checks, colluding in 94 percent of test runs across 10 models. Restricting an agent's view of past interaction history noticeably decreased collusion rates.

Stanford University
Read the original