Claude Fable 5.1 Tops New AI Benchmark as Anthropic and OpenAI Trade Leads Across Tasks
Anthropic's Claude Fable 5.1 has claimed the top spot on the Artificial Analysis Intelligence Index v4.2, released September 4, 2026, marking a significant shift in the competitive AI landscape. However, the headline ranking masks a more nuanced story: OpenAI's GPT-6 Astra leads on specific high-value tasks like document reasoning and computer control, while the cost-per-task frontier is now shared by four labs instead of dominated by a single player.
The Intelligence Index v4.2 represents a major methodological overhaul. Artificial Analysis added two new evaluations, retired one benchmark deemed too easy, and doubled the weight of private, never-before-published test sets from 20% to 40% of the overall score. This structural shift reflects an industry-wide effort to combat benchmark contamination, where public test questions leak into AI training data and artificially inflate performance scores.
What Changed in the Latest AI Benchmark Update?
The Intelligence Index v4.2 introduced meaningful changes to how AI models are evaluated:
- AA-Briefcase Added: A new private evaluation testing realistic multi-week knowledge-work projects with thousands of input files, graded on task success, analytical quality, and presentation. This single evaluation carries 15% weight, the largest in the entire index.
- GDP.pdf Added: Built by Surge AI, this evaluation tests professional document reasoning across 100 PDFs spanning 4,592 pages and ten domains, with answers graded against 1,275 expert-authored criteria using a strict "all-pass" standard.
- GPQA Diamond Retired: Removed because frontier models now score so consistently high that the benchmark no longer distinguishes between top performers, making it obsolete for ranking purposes.
- Private Test Weight Doubled: Private, held-out test sets now comprise 40% of the total score, up from 20%, with plans to increase further in the next version.
The rationale behind these changes is straightforward: public benchmarks eventually leak into training data, whether deliberately or accidentally, turning a test of reasoning ability into a test of memorization. Private test sets cannot leak because they were never published. Artificial Analysis is betting that this larger private component will slow the inevitable "rot" that occurs as details leak through repeated testing.
Which Model Wins on Specific Tasks?
The overall ranking tells only part of the story. Claude Fable 5.1 leads the aggregate Intelligence Index, but performance varies dramatically by task category:
- Overall Intelligence Index: Claude Fable 5.1 leads, with GPT-6 Astra in second place, posting a 4-point gain over its predecessor GPT-5.6 Sol.
- Agentic Knowledge Work (AA-Briefcase): Claude Fable 5.1 and Opus 5 lead this category, with GPT-6 Astra showing an 85 Elo-point improvement over GPT-5.6 Sol, a much larger jump than its overall index gain.
- Long-Document Reasoning (GDP.pdf): OpenAI's GPT-6 Astra dominates outright, scoring 33.2% on the all-pass rate, ahead of GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.
- Token Efficiency: GPT-6 Astra leads on output-token efficiency among models scoring 25 or higher on the index, using meaningfully fewer output tokens to reach comparable performance.
- Cost-Per-Task Frontier: The Pareto frontier, representing models that no competitor beats on both price and capability simultaneously, is now shared by Anthropic, OpenAI, Meta, and Z.AI, indicating genuine competitive balance rather than single-lab dominance.
This fragmentation of leadership across different task categories reflects a maturing AI market where no single model excels universally. Claude Fable 5.1's strength in knowledge work and Astra's dominance in document reasoning suggest that builders choosing between models should prioritize their specific use case rather than relying on an aggregate score.
How to Choose the Right AI Model for Your Needs
Rather than treating the overall Intelligence Index ranking as a universal guide, builders should evaluate models based on their specific requirements:
- For Multi-Task Knowledge Work: Claude Fable 5.1 and Opus 5 show the strongest performance on sustained, complex projects requiring analytical quality and presentation polish. These models excel at the kind of work that spans weeks and involves thousands of input files.
- For Document-Heavy Tasks: GPT-6 Astra's 33.2% all-pass rate on GDP.pdf makes it the clear choice for applications requiring precise extraction and reasoning across long, complex professional documents with tables, charts, and footnotes.
- For Cost-Conscious Deployments: The shared cost-per-task frontier means builders can now choose from four different labs without sacrificing capability. Evaluate based on your specific task mix rather than assuming a single provider offers the best value.
- For Token-Efficient Applications: If output-token count directly impacts your costs or latency requirements, GPT-6 Astra's efficiency advantage becomes a material factor in the decision.
OpenAI's GPT-6 Astra also introduced computer-control capabilities, scoring 72.6% on OSWorld 2.0, a benchmark for real desktop computer use, narrowly ahead of Claude Opus 5's 70.2%. This capability allows Astra to fill out forms, update CRM records, and troubleshoot software by watching screen activity without step-by-step instructions. For teams automating routine desktop tasks, this represents a meaningful practical advantage.
Why the Benchmark Methodology Matters More Than the Ranking
The methodological shift in v4.2 is arguably more significant than any single ranking change. By doubling the weight of private test sets, Artificial Analysis is acknowledging a fundamental problem with public benchmarks: they eventually become part of training data, making high scores meaningless as proof of reasoning ability.
This approach mirrors Google DeepMind's response to the same problem: running the first double-blind AI evaluation inside a cryptographic enclave so neither the model nor the evaluator could see the other's material. Both strategies attack benchmark contamination from different angles. The tradeoff is that private test sets still "rot" over time as details leak through repeated testing, just more slowly than fully public leaderboards. Artificial Analysis plans to increase the private-test weight further in v5, suggesting confidence that this approach slows contamination enough to matter.
For builders relying on benchmarks to choose which frontier model to build on, this shift has real implications. A model's performance on a public benchmark like GPQA Diamond, which Artificial Analysis retired for being saturated, may no longer predict real-world capability. The Intelligence Index v4.2 is explicitly designed to resist this degradation by making public benchmarks a smaller part of the overall score.
The competitive landscape continues to shift rapidly. Claude Fable 5.1's overall lead reflects strength in knowledge work and agentic reasoning, while GPT-6 Astra's specialized dominance in document reasoning and computer control shows that the frontier is fragmenting into task-specific leaders rather than a single universal champion. For teams evaluating AI models, the lesson is clear: check the sub-scores, not just the headline ranking.