Claude Sonnet 5 Just Outperformed Opus on a Major Benchmark. Here's What That Means for AI Work
Claude Sonnet 5, released on June 30, 2026, is Anthropic's mid-tier model that now delivers performance close to its flagship Opus 4.8 on many tasks while costing roughly 40 to 60 percent less. The model introduces a significant shift in how AI teams think about cost versus capability, especially for everyday coding, agentic workflows, and knowledge work that doesn't require the absolute highest performance tier.
What Makes Sonnet 5 Different From Earlier Versions?
Sonnet 5 represents the biggest generation-over-generation leap in the Sonnet line so far, with improvements spanning reasoning, tool use, coding, and knowledge work. The model is specifically built to plan multi-step work, use tools like browsers and terminals, and run autonomously on long tasks, which Anthropic calls its most agentic Sonnet yet.
Three technical changes stand out. First, Sonnet 5 now decides how much to reason on its own rather than requiring manual configuration. Instead of a separate extended-thinking toggle, it adjusts its reasoning depth automatically based on the task in front of it, meaning simple prompts stay fast while harder ones get more deliberation without any setup from you. Second, the model uses a new tokenizer, the same change Anthropic introduced with Opus 4.7, which means the same input can map to roughly 1.0 to 1.35 times more tokens depending on content type. Third, manual extended thinking and non-default sampling parameters like temperature now return errors, so code relying on either will need updates.
How Does Sonnet 5 Actually Perform Against Opus?
Sonnet 5 posts major gains over its predecessor, Sonnet 4.6, across every evaluation Anthropic disclosed, and it lands within a few points of Opus 4.8 on most benchmarks. Two results are particularly striking. On Terminal-Bench 2.1, a test for agentic coding work, Sonnet 5 doesn't just close the gap to Opus 4.8; it passes it, scoring 80.4 percent compared to Opus's 74.6 percent, a jump of more than 13 points over its own predecessor. On GDPval-AA v2, a knowledge-work benchmark, Sonnet 5 edges Opus 4.8 by a slim margin. This marks the first time a Sonnet-class model has outscored the concurrent Opus flagship on any benchmark.
Opus 4.8 still leads on the hardest coding and reasoning tasks, and Anthropic continues to recommend it for cybersecurity work that Sonnet 5 was deliberately restricted from. However, the practical implication is clear: Sonnet 5 handles the broad middle of everyday coding, agentic, and knowledge tasks, and you escalate to Opus only when a workload genuinely needs the last increment of accuracy.
How to Choose Between Sonnet 5 and Opus 4.8 for Your Workload
- Cost and Speed Matter: Sonnet 5 costs $2 per million input tokens and $10 per million output tokens, compared to Opus's $5 and $25 respectively. Choose Sonnet 5 when you're running frequent calls and cost efficiency is a priority across high-volume operations.
- Everyday Professional Work: Sonnet 5 is the default model for Free and Pro users and handles everyday coding, agentic workflows, and knowledge tasks effectively. Use it for most professional workloads unless you hit a performance ceiling.
- Maximum Accuracy Required: Opus 4.8 remains the choice when you need the last increment of accuracy on the hardest coding, most complex reasoning, or cybersecurity work. Escalate to Opus only when Sonnet 5 genuinely can't deliver the precision your task demands.
Sonnet 5 has a 1 million token context window, meaning it can process roughly 100,000 words at once, and supports 128K maximum output. It's available on Claude's Max, Team, and Enterprise plans, as well as through the Claude API and major cloud platforms.
What Changed in Pricing and Availability?
Anthropic made Sonnet 5's pricing permanent in August 2026 at $2 per million input tokens and $10 per million output tokens. The model is now the default for Free and Pro users across the Claude platform. However, there's an important caveat: the new tokenizer means the same input can map to more tokens than before. If you're moving an existing workload over from an earlier model, you should recount a sample of your prompts before assuming your bill stays flat, as real per-task costs can run higher than the sticker price suggests.
Early testers reported that Sonnet 5 finishes complex tasks where previous Sonnet models would stop partway, and that it checks its own output without being asked. For teams that have been running heavier models like Opus because earlier Sonnets couldn't keep up, that assumption is worth revisiting. The capability gap has narrowed significantly while the cost difference remains substantial.
The release of Sonnet 5 fundamentally changes the calculation for AI teams deciding between capability tiers. Rather than a fixed choice between fast-and-cheap or slow-and-powerful, teams now face a cost-versus-accuracy decision at the margin. For most professional and agentic workloads, Sonnet 5 delivers enough performance to justify its lower price, while Opus remains the choice for the hardest problems where that last increment of accuracy matters.