
Science Evaluation Reveals Performance Gap in Autonomous Research Tasks
A benchmark evaluating AI agents on seventy complex scientific problems found that only top systems completed more than half of their assignments without human help. Leading commercial models successfully resolved over sixty percent of tasks, while open source alternatives completed fewer than ten percent.
The Blend
Researchers from Stanford and the open source community, alongside testing organization Artificial Analysis, evaluated leading artificial intelligence systems on automated scientific research tasks. The newly created test suite places software programs inside isolated digital environments with real scientific datasets and software tools, requiring them to complete multi-step projects from start to finish. According to results released by Artificial Analysis, the most capable commercial tools solved over sixty percent of the assigned scientific challenges, while open source alternatives completed under ten percent.
These results highlight both the rapid growth and current limits of automated research tools. For everyday people, advanced artificial intelligence could eventually accelerate drug discovery, climate modeling, and engineering breakthroughs by automating tedious data analysis. However, the evaluation showed that success varies dramatically depending on the subject matter. For instance, top models performed well on mathematical tasks but struggled when handling complex life science problems.
It remains to be seen whether open source projects can narrow this massive performance gap without the vast computing resources of proprietary AI companies. Furthermore, enabling models to spend more processing time on a problem dramatically increases the financial cost per task, forcing users to weigh accuracy against budget constraints. As automated software plays a larger role in laboratory discovery, creating reliable human verification standards will be critical to guarantee that machine-generated scientific findings are accurate and safe.
Written independently by AI News Smoothie from the reporting listed below. Facts belong to the original publishers. Follow the links for their full coverage.
Ingredients
- Artificial Analysis on X: "Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%
Terminal-Bench-Science 0.1 was announced in August 2026, built by @Steve… / X
Artificial Analysis reported that leading commercial AI models completed over 60 percent of autonomous scientific tasks, vastly outperforming open source models which scored ten percent or less.