Logo
FrontierNews.ai

Why Anthropic's Claude Still Dominates AI Spending Even as Open-Weight Models Surge

Open-weight models are processing the majority of AI inference workloads on Vercel's platform, yet Anthropic's Claude models continue to command nearly two-thirds of all spending. This apparent contradiction reveals a fundamental truth about how enterprises choose AI tools: token volume and dollar spend are not the same thing, and consistency in model performance can lock in customer loyalty even as cheaper alternatives proliferate.

What's Driving the Token Volume Shift to Open-Weight Models?

The trend toward open-weight models, which are publicly available and can run on any infrastructure, has accelerated dramatically. In December 2025, open-weight models accounted for just 7% of tokens processed through Vercel's AI Gateway. By August 2026, that share had climbed to 56%, marking the first month these models handled a majority of monthly token volume. The growth has been consistent, rising from 13% in April to 36% in July before jumping to 56% in August.

Chinese-developed models, particularly from providers like DeepSeek, Moonshot AI, and Z.ai, are leading this shift. These models are substantially cheaper to operate than proprietary offerings from US-based AI labs, making them attractive for cost-conscious developers and enterprises. Vercel CEO Guillermo Rauch noted that this milestone likely represents just the beginning, stating that "enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic".

Guillermo Rauch

Why Does Anthropic's Spending Share Remain So High Despite Lower Token Volume?

The answer lies in pricing and performance consistency. While open-weight models handled 56% of Vercel's tokens in August, they accounted for only 14 cents of every dollar spent through the gateway. Anthropic's Claude models, by contrast, captured 64 cents of every dollar, a share that has never fallen below 61% in any month since December 2025.

This gap reflects the economics of AI inference. Open-weight models are cheaper to run because they don't require licensing fees and can be deployed on commodity hardware. Proprietary models like Claude command premium pricing because they deliver higher quality outputs, making them worth the extra cost for tasks where accuracy and reliability matter most. The average price per token across Vercel's gateway fell 23.2% in August alone, marking a third consecutive monthly decline, yet Anthropic maintained its dominant spending position.

How Are Customers Choosing Between Claude Models?

Within Anthropic's own lineup, a significant shift occurred in August. Fable 5, a more capable but expensive model, dropped from 13.2% of total gateway spending in July to just 4.9% in August. Meanwhile, Opus 5, a cheaper alternative, climbed to 22.5% of spending. Vercel's analysis found that 90% of teams using Fable reduced their usage, with most moving those workloads to Opus 5 rather than switching to competitors.

The key insight: Anthropic retained the dollars even as customers migrated to a cheaper model within its own product family. This pattern demonstrates what Vercel's report authors called a critical principle: "Lab loyalty doesn't follow brand, it follows model profile, and consistency wins". When a new model preserves what users valued in its predecessor, the lab retains its customers. When it doesn't, those customers fill the need through other providers.

What Happens When Models Fail to Meet Customer Expectations?

Google's experience with Gemini 3 Flash illustrates the stakes. More than three-quarters of the volume lost by Gemini 3 Flash moved to models from other providers, including OpenAI and Anthropic. The consequence was dramatic: Google's overall share of token volume on Vercel's gateway fell from 30% to 5%, with the Gemini 3 Flash decline alone accounting for 22 of those 25 percentage points.

In contrast, when Z.ai launched GLM-5.3-Flash, the new model was processing three times the daily volume of its predecessor, GLM-5.2, within five days. The difference: GLM-5.3-Flash preserved the capabilities users valued while improving performance and cost efficiency.

How to Navigate the Shifting AI Model Landscape

  • Monitor Token Costs, Not Just Volume: Track both the number of tokens your application processes and the actual dollars spent. A model handling 50% of your tokens might represent only 10% of your spending, or vice versa. Understanding this gap helps identify optimization opportunities and prevents overcommitting to expensive models for tasks that don't require premium quality.
  • Test Model Consistency Across Workloads: Before migrating workloads from one model to another, validate that the new model produces comparable quality outputs for your specific use cases. Anthropic's success shows that customers will stick with a lab if newer models preserve the performance characteristics they depend on, even at lower cost.
  • Evaluate Open-Weight Models for Cost-Sensitive Tasks: As open-weight models now handle 56% of tokens on Vercel's gateway, they've proven viable for many production workloads. Assess whether your application's quality requirements justify proprietary model pricing, or whether open-weight alternatives can deliver acceptable results at significantly lower cost.
  • Plan for Model Agnostic Infrastructure: Vercel CEO Rauch emphasized that enterprise adoption of model-agnostic tools is still early. Building your infrastructure to easily swap models reduces vendor lock-in and positions your application to benefit from future improvements in open-weight or proprietary models without major refactoring.

The data from Vercel's AI Gateway reveals a maturing market where price and performance are increasingly decoupled. Open-weight models are winning on volume and cost, but proprietary models like Claude are winning on spending because they deliver consistent, reliable outputs that enterprises trust for mission-critical tasks. As the market evolves, the winners will be those that maintain consistency in model quality and performance, regardless of whether they're open-weight or proprietary.