1 story in this blend

A benchmark evaluating AI agents on seventy complex scientific problems found that only top systems completed more than half of their assignments without human help. Leading commercial models successfully resolved over sixty percent of tasks, while open source alternatives completed fewer than ten percent.