The Last Frontier: Why AI Still Can't Match Human Intelligence on One Crucial Test
For nearly two years, frontier AI models have failed to match average human performance on SimpleBench, a test designed to measure spatio-temporal reasoning, social intelligence, and resistance to linguistic tricks. That streak just ended. In September 2026, Claude Fable 5.1 scored 86.6% on SimpleBench, surpassing the human baseline of 83.7% for the first time. The milestone marks a symbolic turning point in AI development, but experts say the real challenge to artificial general intelligence (AGI) lies elsewhere.
What Is SimpleBench and Why Does It Matter?
SimpleBench is not your typical AI benchmark. Created in October 2024, it was specifically designed to test abilities that pattern-matching alone cannot solve. Unlike benchmarks that measure raw knowledge or coding skill, SimpleBench evaluates how well systems handle spatial reasoning, understand social dynamics, and resist being tricked by adversarial language. A nine-person unspecialized human baseline scored 83.7%, establishing what an average person could achieve without specialized training.
For nearly two years, every frontier model fell short of this mark. GPT-4, Gemini, and other leading systems all underperformed relative to human participants. That consistency made SimpleBench unique: it was the last public text-only benchmark where average humans reliably outperformed cutting-edge AI. Now that Claude Fable 5.1 has cleared it, no such benchmark remains.
What Does This Tell Us About the Path to AGI?
The clearing of SimpleBench might sound like a major milestone toward artificial general intelligence, but AI researchers caution against reading too much into it. According to Dr. Alan D. Thompson, an AI researcher who tracks progress toward AGI, the remaining gaps to true human-level intelligence are not in text-based reasoning at all. Instead, they lie in three critical areas that text benchmarks cannot measure.
Dr. Alan
Thompson defines AGI as a machine capable of understanding the world as well as, or better than, any human in practically every field, including the ability to interact with the physical world through embodiment. By this definition, current AI systems remain far from AGI despite their impressive text-based performance.
- Embodied Dexterity: The ability to manipulate physical objects with precision and adapt to new environments. All major IQ tests for children under 18 include physical object manipulation tasks, from assembling blocks to handling cards and toys, yet AI systems have no equivalent capability.
- Adaptive Learning: The capacity to learn and adjust behavior based on new experiences in real time, rather than relying on patterns learned during training. Current AI models are static after deployment.
- Sensory Grounding: Integration of multiple human senses, including touch, taste, smell, proprioception, and the ability to detect temperature and texture. Text-based AI lacks access to these fundamental channels of human understanding.
Thompson notes that the median human in 2024-2025 can perform tasks like making a cup of coffee in an unfamiliar kitchen or assembling IKEA furniture. These seemingly simple activities require embodied understanding that no current AI system possesses.
How Should We Interpret Recent AI Breakthroughs?
The same month that Claude Fable 5.1 cleared SimpleBench, OpenAI released GPT-6 Astra to partners, claiming it represents a "generational leap" in AI capabilities. The model achieved extraordinary scores on specialized benchmarks: 96% on GPQA Diamond, a test of graduate-level physics and chemistry; 97.6% on FrontierMath Tier 4 v2; and 100% on ExploitBench, making it the first model to cross OpenAI's "Critical" cyber capability threshold.
Yet even these superhuman performances in narrow domains did not advance the AGI countdown. Thompson's analysis notes that GPT-6 Astra "advances ASI-level capabilities in narrow domains while the intelligence profile remains uneven across task types". Artificial superintelligence (ASI) refers to a machine that performs at the level of an expert human in practically any field, but ASI in narrow domains is not the same as AGI across all domains.
The distinction matters because it reveals a fundamental truth about current AI progress: systems are becoming superhuman in isolated tasks while remaining incomplete as general intelligences. Clearing SimpleBench is a milestone in text-based reasoning, but it does not resolve the embodiment problem that Thompson and other researchers consider essential to true AGI.
What Happens Next in the AGI Race?
With no public text-only benchmark remaining where average humans outperform frontier models, the focus of AI development is shifting. The next generation of progress will likely depend on solving the embodiment, adaptive learning, and sensory grounding challenges that text benchmarks cannot address.
This shift has practical implications for how AI labs allocate resources and how researchers measure progress. Benchmark scores will become less meaningful as a measure of AGI proximity. Instead, real-world performance in robotics, autonomous systems, and adaptive learning environments will become the true test of whether AI is approaching human-level general intelligence.
For now, Claude Fable 5.1's achievement on SimpleBench represents a symbolic endpoint: the moment when AI stopped losing to humans on text-based reasoning tasks. But the race to AGI, according to researchers tracking the frontier, is only entering its most challenging phase.
" }