Why ElevenLabs Thinks Voice, Not Text, Is the Future of AI Interfaces
ElevenLabs, the voice AI startup founded by two Polish childhood friends, is betting that voice will become the dominant interface for artificial intelligence, replacing keyboards and text-based interactions. The company's CEO Mati Staniszewski explained that this vision emerged from a personal frustration: watching foreign films dubbed into Polish with a single monotone narrator, which inspired him and co-founder Piotr to build technology that could deliver expressive, emotionally aware audio across any language.
How Did ElevenLabs Build a Frontier AI Model Without Billions in Funding?
Unlike many AI startups that require hundreds of millions or billions of dollars to build competitive models, ElevenLabs took a different path. The company launched in 2022, a time when audio AI was considered a niche domain with relatively few researchers working on it. This timing proved advantageous because audio models required less computational power than large language models or visual AI systems.
The company's strategy focused on three key advantages. First, they recognized that while audio data is abundant, the real challenge lies in transcribing and annotating it correctly. Second, they assembled a distributed team of top audio researchers by recruiting based on GitHub contributions and published work rather than geographic location. The founders started the company remotely, with team members split between London and Warsaw, hiring talent wherever they could find the best people in the field.
"We started in London. We had a lot of people between London and Warsaw and started a company in a completely remote way. So we wanted to hire the best researchers wherever they were," explained Mati Staniszewski, CEO of ElevenLabs.
Mati Staniszewski, CEO at ElevenLabs
Third, ElevenLabs monetized quickly, generating revenue streams early to fund ongoing model development and research. This approach allowed the company to remain financially healthy while competing with better-funded competitors.
What Products Has ElevenLabs Built to Enable Voice Interfaces?
ElevenLabs began with a text-to-speech model designed to understand context and emotion, then expanded into speech-to-text capabilities. The company has since developed real-time streaming models and conversational AI experiences, all aimed at breaking down language barriers and making voice the primary way people interact with technology.
One practical application gaining traction is voice agents in customer service. Rather than requiring customers to fill out forms or navigate complex menus, voice agents allow people to speak directly to AI systems that understand their needs. This approach has led to richer data collection and improved customer experiences across various sectors, including government services.
Steps to Understand ElevenLabs' Voice-First Strategy
- Text-to-Speech Foundation: ElevenLabs built models that understand emotional context and intonation, moving beyond robotic-sounding narration to create natural, expressive audio that conveys meaning and feeling.
- Real-Time Interaction: The company developed streaming models that enable live conversations with AI, allowing voice agents to respond instantly without the delays that plagued earlier systems.
- Language Barrier Removal: By combining text-to-speech, speech-to-text, and translation capabilities, ElevenLabs aims to let anyone speak to AI in their native language while the system responds in theirs.
- Enterprise Adoption: Companies like MatrixCloud are integrating ElevenLabs' voice AI into contact centers, replacing traditional phone trees and form-based customer service with natural voice interactions.
What Challenges Remain for Voice AI?
Despite rapid progress, ElevenLabs acknowledges significant hurdles. While voice agents perform well in structured customer service scenarios, they still struggle with truly emotional interactions that require nuanced understanding of human sentiment and context. Additionally, while the company can produce quality audio for music and creative applications, it has not yet achieved top-chart quality comparable to professional human artists.
The company has also adopted an unconventional organizational structure to accelerate development. ElevenLabs maintains a flat hierarchy without traditional job titles, encouraging collaboration and rapid decision-making across teams. The company introduced a scoring system to streamline negotiations in sales, allowing technical talent to contribute meaningfully to non-technical discussions.
As voice AI continues to mature, ElevenLabs' vision of voice as the primary interface for AI represents a fundamental shift in how humans will interact with technology. Rather than typing queries into chatbots or clicking through menus, users may soon simply speak to AI systems that understand context, emotion, and intent across any language. This transition could reshape customer service, accessibility for people with disabilities, and the way billions of people interact with artificial intelligence globally.