Indian Startup Sarvam AI Beats OpenAI and ElevenLabs at Understanding Indian Languages
A Bangalore-based startup has quietly outperformed some of the world's largest AI companies at a task that matters enormously in a country with 1.4 billion people and 22 official languages: accurately understanding what people are actually saying. Sarvam AI's latest speech recognition model, called Saaras V3, achieved a word error rate of approximately 19.3% on Indian language benchmarks, beating OpenAI's GPT-4o Transcribe, ElevenLabs' Scribe v2, Google's Gemini 3 Pro, and Deepgram's Nova-3.
Why Does This Matter for Voice AI?
In speech recognition, a lower word error rate means fewer mistakes. But the real story here goes deeper than just a performance metric. India's linguistic landscape is extraordinarily complex. The country's 22 scheduled languages span multiple script families, phonetic systems, and grammatical structures. Add in the widespread practice of code-mixing, where speakers blend Hindi and English mid-sentence, or Tamil and English, or any number of combinations, and you get audio data that would challenge most Western-trained models.
ElevenLabs and other global voice AI companies have built their systems primarily around English and a handful of European languages. When these models encounter the nuances of Indian languages, they often stumble. Sarvam AI took a different approach. The company trained Saaras V3 on over one million hours of multilingual Indian audio, with specific attention to noisy speech environments and code-mixed conversations. This purpose-built strategy appears to have paid off.
The performance gap widened even further when comparing the models on lower-resource Indian languages, suggesting that Saaras V3's architecture and data pipeline were designed specifically for this challenge rather than retrofitted from an English-first approach.
What Capabilities Does Saaras V3 Offer?
Beyond raw accuracy, Saaras V3 includes several features that make it practical for real-world business applications. The model supports native real-time streaming, speaker diarization (the ability to distinguish between different speakers in a conversation), and automatic language detection. These capabilities are particularly suited for call centers, media companies, and enterprise customer service operations that need to function at scale across India's diverse linguistic regions.
Saaras V3 also earned a leading position on the Svarah benchmark, which specifically measures how well models handle Indian-accented English, another critical use case for businesses serving Indian markets.
How Sarvam AI Is Building a Broader India-First Technology Stack
- Speech Recognition: Saaras V3 handles transcription across all 22 scheduled Indian languages plus English, with support for code-mixed speech and noisy environments.
- Text-to-Speech: Bulbul, Sarvam's text-to-speech system, complements the speech recognition capabilities and allows businesses to generate natural-sounding audio in Indian languages.
- Translation and Conversation: The company offers translation services and a conversational agent platform called Sarvam Samvaad, creating an integrated ecosystem for multilingual AI applications.
Saaras V3 sits within this broader India-first technology stack, which covers all 22 scheduled Indian languages plus English. This comprehensive approach suggests that Sarvam AI is positioning itself not just as a speech recognition vendor, but as a foundational AI infrastructure provider for the Indian market.
The company is one of twelve startups collaborating with the Indian government under the IndiaAI mission, a national initiative aimed at developing indigenous multilingual and multimodal AI technologies. This government backing signals both the strategic importance of localized AI development and the recognition that global AI companies may not adequately serve India's unique linguistic needs.
What Does This Mean for the Global Voice AI Market?
The emergence of Sarvam AI as a leader in Indian language voice recognition highlights a broader trend in AI development. While companies like ElevenLabs have built impressive voice technology, their models are optimized for English and Western languages. As AI adoption spreads globally, the limitations of one-size-fits-all approaches become increasingly apparent. Regions with diverse, lower-resource languages, complex linguistic patterns, and high code-mixing rates require purpose-built solutions.
Sarvam AI's success suggests that the future of voice AI may not be dominated entirely by a handful of global players. Instead, regional specialists that understand local language nuances, cultural contexts, and business needs may carve out significant market share. For enterprises operating in India, this development offers an alternative to relying solely on global voice AI platforms that may not perform as well on Indian languages.