The 37-Point Gap: How OpenAI's AGI Claim Reveals the Benchmark Problem Nobody's Talking About
OpenAI's declaration that artificial general intelligence (AGI) has arrived rests on benchmark scores that differ dramatically depending on who's measuring. The company's new Astra model, released as GPT-6, scored 99.9% on ARC-AGI-3 using OpenAI's internal evaluation setup, while the benchmark's creators measured the same model at 62.7% using their reference protocol. Both numbers come from real evaluations, but they tell very different stories about whether AI has crossed a historic threshold.
The announcement came swiftly. On September 6, Jensen Huang, CEO of Nvidia, posted a single sentence on X: "From ChatGPT to o1 to Astra in 4 years, AGI has arrived. Congratulations @OpenAI team." Three days earlier, OpenAI President Greg Brockman had used nearly identical language with the press: "Welcome to the AGI era." The declarations were received largely as confirmation of a milestone, but the gap between how the model performed and how that performance was framed raises a fundamental question about how the AI industry measures progress.
Why Do the Same Benchmarks Produce Such Different Scores?
The discrepancy stems from different evaluation protocols. OpenAI's internal harness and the benchmark creators' reference harness are not the same system, and the same model produces different results under each. The ARC-AGI-3 organization, which built the test, published 62.7% as the score from their standard evaluation setup, the one they consider valid for comparing models across the field. OpenAI's internal setup produced 99.9%. That 37-percentage-point gap is the difference between a strong benchmark result and a declaration of arrival.
To understand why this matters, it helps to know what ARC-AGI-3 actually tests. The benchmark was designed to resist the optimization tricks that inflated earlier AI benchmarks. Previous tests were "saturated," meaning models had absorbed training data resembling the test questions and pushed numbers up without actually improving at novel reasoning. ARC-AGI-3 changed the question design to require abstraction over pattern retrieval, making it harder to game. A score of 62.7% on that test, under the reference harness, is genuinely strong and far above anything that existed two years ago.
But here's the structural issue: when a company measures its own product against a definition it wrote, using a protocol it designed, the result is different from one measured by an external standard. OpenAI has an internal definition of AGI. By that definition, Astra may qualify. The problem is that no stable, agreed definition of AGI exists across the field. Researchers who raised objections about the announcement were precise about this distinction. Their target was not the model's capability, but the jump from performance to declaration.
What Does Independent Verification Show?
As of the time these announcements were made, Astra did not appear in Artificial Analysis Intelligence Index or Arena.ai's rankings, the two most-watched independent evaluation systems for frontier models. Independent verification for major frontier releases typically arrives within days. The absence here is not evidence of poor performance; it means the external validation that would normally accompany a claim of this magnitude had not materialized yet.
The 62.7% score from the benchmark creators is not a refutation of Astra's capabilities. It represents a genuinely strong result on a test designed to resist the optimization tricks that have inflated earlier benchmarks. The issue is semantic and structural: "Scores well on these benchmarks" and "AGI has arrived" require a step in between, a stable, agreed definition of AGI against which arrival can be measured. No such definition exists across the field.
How to Interpret AI Benchmark Claims Like These
- Check the Protocol: Ask whether the score comes from the benchmark creator's reference harness or from the company's internal evaluation setup. Different protocols can produce dramatically different results on the same model.
- Look for Independent Verification: Established independent evaluation systems like Artificial Analysis Intelligence Index and Arena.ai typically verify major frontier model claims within days. The absence of such verification is a signal to wait for external confirmation.
- Distinguish Performance from Declaration: A model performing well on a benchmark is different from a claim that a historic threshold has been crossed. Strong performance requires a stable, agreed definition of what the threshold means.
- Consider the Incentive Structure: When a CEO publicly congratulates a customer for clearing a historic threshold, the announcement is structurally a commercial claim. The chip market for AI infrastructure moves on narrative, and framing shapes investment decisions.
The vocabulary of AI is evolving rapidly, and so are the terms used to describe how models work. Concepts like "opaque recurrence," the reasoning technique in OpenAI's Astra model, have emerged recently and have AI safety researchers concerned. But the language around AGI itself remains nebulous. OpenAI CEO Sam Altman once described AGI as the "equivalent of a median human that you could hire as a co-worker," while OpenAI's charter defines it as "highly autonomous systems that outperform humans at most economically valuable work." Google DeepMind's understanding differs slightly, viewing AGI as "AI that's at least as capable as humans at most cognitive tasks." Even experts at the forefront of AI research are confused about what AGI actually means.
The reasoning models that power systems like o1 and Astra represent a real shift in how AI systems approach problems. These models use chain-of-thought reasoning, breaking down problems into smaller, intermediate steps to improve the quality of the end result. It usually takes longer to get an answer, but the answer is more likely to be correct, especially in logic or coding contexts. Reasoning models are developed from traditional large language models (LLMs) and optimized for chain-of-thought thinking through reinforcement learning (RLHF), a technique that trains AI systems based on feedback.
The four-year progression Huang named, from ChatGPT to o1 to Astra, is real as a trajectory of capability. Each model in the sequence extended what was possible. But framing that sequence as a journey with an endpoint, and announcing the endpoint, makes a different claim: it says arrival is distinguishable from the journey, that there is a line rather than a slope, and that Astra crossed it. The benchmark number behind that claim was 62.7%. The people who built the test published it. The question it raises, whether the AGI declaration is a measurement or a market move, was there from the start. The coverage largely answered it without asking.
" }