From Voice AI to Deepfake Detection: Why Resemble AI Pivoted Its Entire Business
Resemble AI, a company that built its reputation creating synthetic voices for gaming studios, has completely pivoted its business model to focus on detecting deepfakes across all media formats. The shift reflects a stark reality: as audio-visual AI (multimodal AI) becomes more sophisticated, the technology that creates convincing synthetic content is increasingly weaponized for fraud, identity theft, and non-consensual intimate imagery. The company's H1 2026 Deepfake Threat Report found at least 15,736 people victimized across 821 documented attacks in just six months, with roughly 3.46 million malicious files involved.
What Prompted a Voice AI Company to Abandon Its Core Product?
Resemble AI was founded in 2019 to solve a problem in the gaming industry: voice actors couldn't keep pace with how quickly studios shipped games, and human talent didn't scale. The company built generative voice technology by going deep into model architecture, training data, and the mechanics of how synthetic speech is actually produced. That expertise turned out to be unexpectedly valuable, but not in the way founders Zohaib Ahmed and Saqib anticipated.
"Understanding how voice AI is built at the model level gives us a major advantage at identifying clones and gen AI voices in the wild," said Zohaib Ahmed, Co-Founder and CEO at Resemble AI.
Zohaib Ahmed, Co-Founder and CEO at Resemble AI
The company's deep knowledge of how synthetic voices are constructed became the foundation for detection work. Early on, Resemble AI researchers suspected that as voice AI improved, it would increasingly be misused. They began building detection tools alongside their generative models, releasing open-source projects like Resemblyzer, a speaker verification tool that can match any piece of audio to an enrolled identity.
But the threat landscape shifted faster than anticipated. Deepfake incidents multiplied across all media types, not just audio. The H1 2026 report revealed a particularly disturbing trend: one in six of the documented attacks involved non-consensual intimate imagery of adults or children. That finding forced a strategic reckoning. The company decided to commit entirely to detection work, retiring its commercial voice generation business and putting the full force of the organization behind multimodal deepfake detection.
How Does Resemble AI's New Detection System Actually Work?
Resemble AI now offers five core detection APIs (Application Programming Interfaces, or tools that software developers use to access specific functions) designed to create what the company calls an "enterprise trust stack." Each API addresses a different question about media authenticity and origin.
- Detect: Determines whether audio, video, or image content is real or synthetic, delivering a verdict in under 300 milliseconds. The system runs on DETECT-World, which checks whether media is physically consistent rather than matching signatures from known generators, meaning it can identify deepfakes created by models released after the detection system was trained.
- Watermarker: Embeds an imperceptible watermark at the point of generation and reads it back even after compression, editing, and re-encoding. This mechanism is compatible with C2PA (Coalition for Content Provenance and Authenticity) standards and aligns with EU AI Act Article 50 disclosure requirements.
- Identity: Enrolls a known voice or likeness, then matches incoming media against that enrolled identity to verify whether the person on a call or in a video is actually who they claim to be.
- Signal: Performs fraud pattern matching against preset and custom libraries, similar to how Shazam identifies songs. Importantly, it does not transcribe content, store media, or expose personally identifiable information.
- Intelligence: Converts detection scores into plain-language reasoning and structured data formats, allowing analysts, regulators, or courts to understand and act on results rather than accepting them on faith.
The company's DETECT-World audio deepfake detector currently achieves 99.5% detection accuracy on public benchmarks, making it the most accurate audio deepfake detector available by those measures. This performance reflects years of research into how synthetic speech is actually constructed at the model level.
Why This Pivot Matters for the Broader AI Industry
Resemble AI's decision to abandon a profitable product line signals a larger truth about multimodal AI development: the same capabilities that enable creative applications like voice acting and video generation can be weaponized at scale. The company's open-source voice models have been downloaded more than 17 million times on Hugging Face, a platform where researchers share AI models, proving that foundational detection research can reach massive audiences and withstand public scrutiny.
The company is not selling voice generation to new customers and has retired more than 300 older blog posts about voice cloning, text-to-speech, and audio editing. Existing customers are still supported, but the company will not take on new voice AI clients. Instead, Resemble AI is laser-focused on one problem: building the most accurate multimodal deepfake detection tools possible.
The threat data backing this decision is sobering. In just six months of 2026, documented deepfake attacks victimized over 15,700 people, with roughly 3.46 million malicious files involved. The majority of those files were image and video, not audio, underscoring why the company expanded detection capabilities beyond its original audio expertise.
Steps to Understanding Deepfake Detection in Your Organization
- Assess Your Risk Profile: Determine which media types (audio, video, image) pose the greatest risk to your organization. Financial institutions, government agencies, and media companies face different threat landscapes.
- Evaluate Detection Accuracy: Look for detection systems that report performance on public benchmarks and can identify deepfakes created by models released after the system was trained, not just known generators.
- Consider Watermarking at Generation: If your organization generates synthetic media, embedding imperceptible watermarks at creation time provides a permanent chain of custody that survives editing and re-encoding.
- Plan for Identity Verification: For high-stakes communications like financial transactions or sensitive meetings, implement identity verification systems that match incoming media against enrolled identities.
- Require Explainability: Ensure any detection system can explain its verdict in plain language, not just return a confidence score, so analysts and decision-makers can understand and act on results.
Resemble AI's pivot reflects a maturation moment in the audio-visual AI industry. As multimodal models become more capable, the detection and authentication infrastructure must evolve in parallel. The company's decision to abandon its original business model in favor of detection work suggests that, at least for this founder team, the most valuable application of their expertise is not in creating synthetic media, but in verifying what is real.