GPT-6 Astra solves interactive intelligence benchmark with human beating efficiency
Models & ResearchSuperintelligence · 2h ago · also in Superintelligence

GPT-6 Astra solves interactive intelligence benchmark with human beating efficiency

Evaluators verified that GPT-6 Astra reached a 99.9 percent success rate on the ARC-AGI-3 puzzle environment. The system required fewer actions per level than human baselines while managing complex internal notation systems to track game states. Achieving these results required thousands of dollars in computational spend via a custom testing harness.

ARC PrizeOpenAI

The Blend

OpenAI's latest model, GPT-6 Astra, achieved near-perfect scores on the ARC-AGI-3 benchmark, an evaluation designed to test how well artificial intelligence can navigate unfamiliar, interactive puzzle environments. As detailed on the official ARC Prize site, the model solved 99.9 percent of the benchmark's abstract games when paired with a specialized testing system. Remarkably, the AI completed 96 percent of the puzzle levels using fewer moves than the average human participant.

To achieve these results, the system did not just guess at random. Instead, it invented its own shorthand notation of rules and symbols to map out each game's mechanics, track its progress, and plan future actions. For everyday users, this signals a shift from AI models that merely generate text to automated agents capable of independently reasoning through complex, multi-step problems without explicit instructions.

However, this level of problem solving comes with a steep price tag. Running the benchmark tests required nearly 19,000 dollars in computing expenditure for the model to work through the puzzles. While human players cost more in terms of hourly wages during testing, the actual electricity used by a human brain costs less than a penny per game, highlighting the massive energy gap that still exists between machine intelligence and biological thinking.

It remains unclear whether these high-cost reasoning techniques will translate affordably into consumer software or remain locked behind enterprise budgets. If creating a compact mental model of a problem requires thousands of dollars in cloud compute, developers may need to discover new ways to streamline internal logic before autonomous software agents become practical for everyday tasks.

Written independently by AI News Smoothie from the reporting listed below. Facts belong to the original publishers. Follow the links for their full coverage.

Ingredients

Read the original