
Models & ResearchSuperintelligence · 3h ago
Discrepancy in AI Test Scores Highlights Benchmark Evaluation Issues
OpenAI introduced its flagship GPT-6 Astra model alongside claims of peak performance across multiple reasoning evaluations. The ARC Prize Foundation later noted that Astra achieved 98.6 percent on its benchmark using a custom OpenAI adapter, but scored 62.7 percent under the standard test environment. The result demonstrates how software harnesses and evaluation settings heavily influence reported artificial intelligence capabilities.
OpenAIARC Prize FoundationGreg Brockman
Read the original