Logo
FrontierNews.ai

Cartesia's Sonic-3.6 Dethrones ElevenLabs in Text-to-Speech Quality Rankings

Cartesia has released Sonic-3.6, a real-time text-to-speech model that now holds the top position on both Artificial Analysis speech leaderboards, surpassing ElevenLabs' widely-used Eleven v3. The new model scores 1,283 Elo on the Provider Voice board and 1,123 on the Controlled Voice board, where every model is tested using the same eight reference voices to isolate pure synthesis quality from voice catalog differences.

What Makes Sonic-3.6 Different From Other Text-to-Speech Models?

Sonic-3.6 uses state space models rather than transformer architecture, a technical choice that Cartesia argues allows it to break the traditional tradeoff between speed and naturalness. The model delivers sub-90-millisecond time-to-first-audio, meaning it begins speaking in under a tenth of a second. This speed matters for real-time applications like customer service agents and voice interfaces where users expect immediate responses.

The model includes features built specifically for conversational AI rather than audiobook narration. Users can embed non-verbal expressions like [laughter] directly into transcripts, clone voices from roughly 10 seconds of audio, and use custom pronunciation dictionaries with phonetic overrides for tricky words like "subpoena." The API also exposes controls for speed, volume, and emotion parameters, plus native handling of alphanumerics so order numbers and confirmation codes read correctly without preprocessing.

How Does Sonic-3.6 Compare on Price and Performance?

Pricing represents a significant competitive advantage. Artificial Analysis normalizes Sonic-3.6 at $49.00 per million characters, exactly half the cost of ElevenLabs Eleven v3 at $100.00 per million characters. For context, Speechify's Simba 3.2 costs $10.00 per million characters but scores lower on quality benchmarks at 1,240 Elo.

Cartesia's pricing structure uses monthly credit tiers rather than per-character billing. The Scale tier at $299 per month includes approximately 10,667 text-to-speech minutes and 15 concurrent requests. For voice agents handling live calls, the company charges separately at $0.06 per minute.

What Are the Key Capabilities for Developers and Enterprises?

  • Deployment Model: Sonic-3.6 is available as a hosted API in beta, not as self-hosted weights, meaning developers rent access rather than running the model locally on their own servers.
  • Target Markets: Cartesia positions Sonic for financial services, healthcare, retail and e-commerce, logistics, recruiting, SaaS support, consumer companion apps, and media localization across multiple languages.
  • Use Cases: The model powers inbound support agents, outbound qualification calls, IVR replacement systems, appointment reminders, sales-training simulators, audio localization, and in-product voice user interfaces.
  • Pricing Tiers: Solo developers and startups access free and $5 Pro tiers, while scaleups running contact centers choose Startup ($49/month) or Scale ($299/month) plans, with regulated enterprises receiving custom pricing for data protection agreements and single sign-on.

Why Does the Controlled Voice Board Matter More Than Provider Voice?

The Controlled Voice board isolates the synthesis engine from the voice catalog by cloning every model onto identical reference voices. This methodology reveals whether Sonic-3.6's victory comes from better underlying technology or simply a superior voice library. Sonic-3.6 leads the Controlled board with 1,123 Elo, followed by Sonic-3.5 in second place and ElevenLabs Eleven v3 in third, confirming that the engine itself improved rather than just the available voices.

Cartesia's launch demonstrations showcase English with natural pauses and filler words, plus Hinglish code-switching between Hindi and English, suggesting the model handles multilingual and mixed-language scenarios that many real-world applications require.

What Should Developers Know About Latency Claims?

Cartesia states sub-90-millisecond time-to-first-audio for Sonic-3.6 and 100-millisecond transcript latency for its Ink-2 speech-to-text model. However, these figures represent vendor-stated model latency in isolation, not measured end-to-end round trips across networks and integrations. Developers implementing Sonic should benchmark their own real-world latency before committing to production deployments.

The model remains in beta on Cartesia's API, with documentation still listing Sonic 3.5 as the stable version. Partners continue offering Sonic 3.5 until the newer version exits beta, so adoption timelines may vary across platforms and integrations.