Logo
FrontierNews.ai

Claude Fable 5.1 and Claude Mythos 5.1 Face New Competition as Sakana's Multi-Model System Outperforms Both

Sakana AI has released a new multi-model orchestration system called Fugu Ultra v2 that surpassed Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra on several benchmark tests, challenging the industry assumption that bigger single models always perform better. The Tokyo-based company also introduced Fugu Max, a more cost-effective variant designed to balance performance with affordability.

How Does Sakana's Multi-Model Approach Work Differently?

Rather than relying on one large language model (LLM), which is a type of artificial intelligence trained on vast amounts of text to understand and generate human language, Sakana's Fugu system assigns different tasks to different AI models based on what each one does best. This orchestration strategy allows the system to achieve higher overall performance than any single model could deliver alone. The approach represents a philosophical shift in how companies are thinking about AI capability.

Fugu Ultra v2 demonstrated measurable advantages across multiple evaluation benchmarks. Notably, it outperformed Claude Fable 5.1 and GPT-6 Astra in several tests, even though neither of those models is included in Fugu Ultra v2's own model ensemble. This suggests that intelligent task routing and model selection can sometimes exceed the capabilities of individual frontier models.

What Are the Key Differences Between Fugu Ultra v2 and Fugu Max?

Sakana released two versions to serve different use cases and budgets. Here's how they compare:

  • Performance Focus: Fugu Ultra v2 is optimized for maximum capability and achieved the highest scores on benchmark tests, making it suitable for complex reasoning tasks and applications where accuracy is paramount.
  • Cost Efficiency: Fugu Max supports a wider variety of compatible models and achieved the highest performance-to-cost ratio on Terminal Bench 2.1, an agent performance evaluation, while keeping expenses down.
  • Pricing Structure: Fugu Ultra v2 costs $5 per million input tokens, while Fugu Max costs $2 per million input tokens, making the latter roughly 60% cheaper for basic input processing.

Both systems are available through an application programming interface (API), which is a technical interface that allows software to communicate with other software, on a pay-as-you-go basis or through monthly subscription plans ranging from $20 to $200.

Why Does This Challenge the Current AI Industry Narrative?

For the past several years, the dominant strategy in AI development has been to build larger and larger single models, with companies competing on parameter count and raw computational power. Claude Opus 5, Claude Fable 5, and GPT-6 Astra all represent this "bigger is better" philosophy. Sakana's results suggest that intelligent orchestration of multiple models, including smaller and more specialized ones, can outperform this approach on certain benchmarks.

This finding aligns with broader industry observations that model size alone does not guarantee superior performance. The ability to route tasks intelligently, combine different models' strengths, and optimize for specific problem types may matter as much as raw model scale. For organizations evaluating AI tools, this opens up new possibilities for achieving high performance without necessarily adopting the largest available models.

What Does This Mean for Anthropic's Claude Lineup?

Anthropic's Claude family includes Claude Fable 5, Claude Fable 5.1, and Claude Mythos 5.1, which the company has positioned as high-performance options with up to 45% lower pricing than earlier versions. The emergence of Sakana's competing system that outperforms Claude Fable 5.1 on benchmarks suggests that the competitive landscape for AI models is becoming more fragmented. Organizations can no longer assume that a single model from a major vendor will be optimal for all use cases.

The practical implication is that enterprises and developers now have more options for achieving their performance targets. Rather than defaulting to the largest or most expensive model available, teams can evaluate whether a multi-model orchestration approach might deliver better results at lower cost. This shift could accelerate adoption of hybrid AI strategies where different models handle different tasks within the same application.

" }