Logo
FrontierNews.ai

ElevenLabs Falls Behind on Voice Cloning Quality, New Benchmark Shows

ElevenLabs, long considered the default choice for voice cloning, is losing ground to competitors on a critical quality metric: how well cloned voices actually sound like the original speaker. A new benchmark released by Hume in September 2026 tested 11 voice cloning models and found that ElevenLabs' flagship Eleven v3 scored 2.91 out of 5 on speaker similarity, ranking last among the vendors tested.

The findings challenge ElevenLabs' market dominance at a moment when voice cloning is becoming essential infrastructure for AI applications, from customer service agents to content creation. The benchmark tested each model on 25 different reference voices with seven different prompts, with three blind raters scoring each output on how closely it matched the original speaker.

Which Voice Cloning Models Performed Best?

Fish Audio's S2-pro model led the field with a score of 4.03, followed by Cartesia's sonic-3.5 at 3.70. ElevenLabs' Multilingual v2 scored 3.68, placing it third, but the company's newer Eleven v3 model fell significantly behind. The gap between the top performer and ElevenLabs' flagship is substantial: Fish Audio scored 39% higher than Eleven v3 on the same test.

The benchmark also revealed a crucial insight about voice cloning tradeoffs. Cartesia's sonic-3.6-beta topped the naturalness category at 4.36 out of 5, yet ranked eighth on speaker identity. Inworld's TTS-2 achieved the highest audio quality score at 4.61 but didn't lead on similarity. This suggests that vendors are optimizing for different qualities, and the "best" choice depends on what matters most for a specific use case.

How Do Voice Cloning Platforms Compare on Practical Requirements?

Beyond raw quality scores, voice cloning platforms differ significantly in how much audio they need, how they verify consent, and what they charge. Understanding these differences is critical for teams choosing a platform:

  • Audio Requirements: ElevenLabs recommends 1 to 2 minutes of clean audio for instant cloning but requires at least 30 minutes for professional cloning. Cartesia needs just 10 seconds for instant clones, though using up to 60 seconds improves accent preservation. Inworld works from as little as 3 seconds, while Fish Audio requires about 10 seconds for instant clones.
  • Consent Verification: Only three platforms document technical consent checks for professional clones: ElevenLabs uses Voice Captcha, Fish Audio performs live ownership verification, and Resemble AI requires explicit, verifiable consent. The others rely on terms of service and attestation.
  • Pricing Per Million Characters: Inworld's TTS-2 Flash falls to $7 per million characters on higher plans, making it the cheapest option. Fish Audio charges a flat $15 per million UTF-8 bytes. ElevenLabs v3 costs $100 per million characters, among the highest in the market.

ElevenLabs does offer the strictest consent gate for professional clones: the voice owner must read on-screen text aloud through Voice Captcha, a security measure that other platforms don't require. This appeals to enterprises concerned about deepfake risks, but it comes at a cost in both setup friction and pricing.

Why Does This Matter for Teams Building Voice AI?

The benchmark results suggest that teams should no longer assume ElevenLabs is the automatic choice. For applications where speaker similarity is critical, Fish Audio and Cartesia now offer measurably better quality. For cost-sensitive deployments at scale, Inworld and Fish Audio provide significantly cheaper alternatives. For teams in regulated industries where consent verification is non-negotiable, ElevenLabs' professional cloning path remains the most rigorous option.

The voice cloning market is also expanding in breadth. Inworld supports over 200 languages and locales, while Gradium covers only five European languages. Hume's Octave supports 11 languages and emphasizes emotion-directed delivery. This fragmentation means the best platform depends entirely on what a team needs to build.

ElevenLabs' position as the category default is being tested. The company still dominates in brand recognition and has the most mature consent infrastructure for professional use. But on the metric that matters most to users, speaker similarity, it is now trailing competitors that were less well-known just months ago. For teams evaluating voice cloning platforms in late 2026, the benchmark provides clear evidence that the market has matured beyond a single dominant player.