ElevenLabs Faces New Pressure as Microsoft Slashes Transcription Prices to 10 Cents Per Hour
Microsoft has released MAI-Transcribe-2, a speech recognition model priced at just 10 cents per hour of audio, undercutting ElevenLabs, OpenAI, and Google while claiming superior speed and accuracy. The move represents a dramatic 72% price cut from Microsoft's previous transcription model released just five months ago, potentially reshaping the competitive landscape for voice AI technology.
What Changed in Microsoft's Latest Transcription Model?
MAI-Transcribe-2 launched on September 3rd with significant technical improvements over its predecessors. The model now supports 60 languages, up from 43 languages in June's MAI-Transcribe-1.5 release, and includes features that competitors typically charge premium prices for as add-ons.
Microsoft bundled several enterprise-focused capabilities into the base product that previously required separate purchases or higher-tier subscriptions:
- Speaker Diarization: Automatically identifies who said what in multi-person recordings, transforming raw audio into organized, attributed transcripts rather than undifferentiated text.
- Word-Level Timestamps: Attaches precise time markers to every word, enabling search, editing, and synchronization with video content for accessibility and compliance purposes.
- Keyword Biasing: Allows developers to feed the model domain-specific terminology, product codes, and employee names so it stops misinterpreting specialized jargon.
- Configurable Output Styles: Offers a "verbatim" mode preserving every filler word and false start for legal teams, alongside a "clean" mode that strips fillers for readable captions and notes.
- Code Switching: Handles conversations that drift between languages mid-sentence, explicitly supporting language pairs like Hinglish and Spanglish common in customer service environments.
- Automatic Language Identification: Detects the language being spoken without requiring users to declare it in advance.
The model ranks first on the FLEURS benchmark across 60 languages with an average word error rate of 5.2%, meaning roughly one word in 20 is incorrect. On the Artificial Analysis leaderboard, which tests models through their public APIs to measure real-world performance, MAI-Transcribe-2 ranks second and defines the accuracy-latency Pareto frontier, meaning no competitor beats it on accuracy without sacrificing speed, and none beats it on speed without losing accuracy.
How Does the Pricing Strategy Reshape the Market?
The 10-cent-per-hour price point represents a seismic shift in transcription economics. For an enterprise processing 100,000 hours of call-center audio annually, a modest volume for a large bank or telecom company, the annual bill drops from $36,000 under the previous pricing to $10,000. At that cost level, transcription stops being a line item that finance departments scrutinize and becomes a commodity service embedded into workflows.
Microsoft achieved this pricing through architectural efficiency. The company built MAI-Transcribe-2 to run at substantially lower computational cost than competing models, processing audio up to 10 times faster than OpenAI's GPT-Transcribe, seven times faster than ElevenLabs' Scribe v2, and five times faster than Google's Gemini 3.5 Transcribe. In batch transcription scenarios where throughput matters more than latency, this efficiency translates directly to lower infrastructure costs, allowing Microsoft to pass savings to customers.
The release cadence itself signals Microsoft's strategic intent. The company shipped three transcription models in five months: MAI-Transcribe-1 in April at $0.36 per hour, MAI-Transcribe-1.5 in June, and MAI-Transcribe-2 in September at $0.10 per hour. Each release expanded language coverage by roughly 40% while adding features competitors gate behind premium tiers. This pattern reflects a team that has stabilized its core architecture and is now scaling data and model size, the phase where speech models typically improve quickly and predictably.
Why Is ElevenLabs Responding With Leadership Changes?
ElevenLabs appointed Ashley Kramer as Chief Revenue Officer on September 2nd, just one day before Microsoft's announcement, signaling the company's commitment to defending its enterprise position. Kramer joins with a mandate to translate ElevenLabs' research and product capabilities into sustained enterprise growth, focusing on strengthening the company's enterprise business, expanding customer and partner relationships, and developing commercial strategy for the next growth phase.
ElevenLabs currently serves customers across 80 countries and reports reaching $600 million in annual recurring revenue, with its technology used within 69% of Fortune 500 companies. Major customers include Meta, Klarna, Stripe, Hasbro, and Deutsche Telekom. However, the company's competitive position has shifted. In June, Artificial Analysis ranked ElevenLabs' Scribe v2 third on its word-error-rate leaderboard, behind Alibaba's Fun-Realtime-ASR-preview and ahead of other competitors. Microsoft's climb to second place suggests it has now cleared ElevenLabs on the metrics that matter most to enterprise buyers.
"Customers value ElevenLabs for creating AI experiences that people enjoy interacting with," Kramer stated, noting that conversations with customers and partners during her first week reinforced this perspective.
Ashley Kramer, Chief Revenue Officer at ElevenLabs
Kramer's appointment reflects ElevenLabs' strategic pivot toward human-centered interaction as its competitive differentiator. Rather than competing purely on transcription accuracy and price, the company is positioning its technology around natural voice and audio interaction designed to make AI-generated communication more expressive and human-like. ElevenLabs' portfolio spans AI voice models, speech, sound and music, alongside tools designed to support AI agents and creative workflows, with technology positioned across use cases including sales, customer support, operations, marketing, and media production.
How to Evaluate Transcription Models for Enterprise Use
- Benchmark Context: Understand what each benchmark actually measures. FLEURS tests read speech from native speakers on identical content across languages, making it ideal for comparing multilingual performance but not representative of real-world noisy audio. Artificial Analysis tests models through their public APIs using simulated agent conversations, European Parliament speeches, and corporate earnings calls, weighting heavily toward English business speech, making it more representative of actual enterprise use cases.
- Real-World Performance Metrics: Request per-language word-error-rate breakdowns rather than averages, since averaging across 60 languages includes low-resource languages where every model struggles. Ask vendors for performance data on your specific audio conditions, whether that means background noise, overlapping speech, or domain-specific terminology.
- Total Cost of Ownership: Calculate not just per-hour pricing but infrastructure costs. A model running at 300 times real-time speed requires a fraction of the GPU-hours of one running at 30 times real-time, potentially offsetting higher per-hour fees. Factor in whether critical features like speaker diarization and keyword biasing are included in base pricing or require premium tiers.
- Feature Bundling: Evaluate which capabilities matter for your use case. Legal and compliance teams need verbatim transcription with timestamps. Customer service operations need speaker diarization and code-switching support. Content creators need clean output suitable for captions. Determine whether you need all features or can optimize for your specific workflow.
Microsoft's strategy of building frontier-class models one modality at a time, then swapping them into products that previously relied on OpenAI's technology, represents a significant shift in how the world's most valuable software company intends to compete in AI. Transcription is the modality where this plan has moved fastest, and MAI-Transcribe-2 is its clearest proof point yet. The company hired Mustafa Suleyman from Inflection AI in March 2024 along with most of Inflection's staff, a move that Salesforce CEO Marc Benioff interpreted as a declaration of intent to build independent frontier models rather than depend on OpenAI long-term.
For ElevenLabs and other voice AI competitors, the challenge is clear. Microsoft has demonstrated it can build commodity-grade transcription at scale and price, forcing rivals to either compete on price, differentiate on features and user experience, or focus on adjacent markets where transcription is a component rather than the primary product. ElevenLabs' appointment of Kramer and its emphasis on human-centered interaction suggest the company is choosing the differentiation path, betting that enterprises value natural-sounding voice and expressive audio interaction enough to justify premium pricing over Microsoft's commodity offering.