How a Chinese AI Startup Is Challenging Google's Video Generation Dominance
HiDream.ai, a Chinese startup founded in 2023, has developed video and world-model technology that competes with Google DeepMind's offerings while operating at significantly lower cost. The company's HiDream-O1-World model scored 80.9 on WBench's Navi leaderboard and ranked first when it launched in August 2026, though newer models have since moved ahead. What sets HiDream.ai apart is not just technical capability, but how it achieves that capability: the company has reduced video training costs to roughly one-fifth of what it considers the industry average, according to founder Tao Mei.
What Makes HiDream.ai's Approach Different From Tech Giants?
HiDream.ai's strategy reflects a deliberate choice to compete on efficiency rather than raw scale. Mei, who spent 12 years at Microsoft Research Asia before leading computer-vision research at JD.com, explicitly avoided a head-on race with large language models (LLMs), which are AI systems trained on vast amounts of text data and require thousands of specialized computing processors. Instead, he focused on image and video generation, where the computational demands were substantial but more manageable.
The company developed a "dual-model" approach: it tested ideas using cheaper image models before carrying them into video development. This reduced overall training costs significantly. Mei told The Paper that HiDream.ai's first video model, released in August 2023, was "terrible," but the company learned from that failure and shifted toward a Diffusion Transformer architecture, a type of AI system that generates images or video by gradually refining random noise into coherent content.
HiDream-O1-World uses what the company calls a Unified Transformer architecture to process text, images, video, and spatial information within a single framework. Crucially, it separates scene geometry from visual appearance, which helps preserve spatial consistency when users navigate and interact with generated environments. This matters because interactive world models, which let users navigate and modify generated scenes in real time, represent the next frontier in AI content generation.
How Is HiDream.ai Building a Business Around AI Models?
The harder lesson for HiDream.ai came from customer feedback. The company initially assumed that better model capability would itself be a compelling product. But enterprise customers wanted complete software solutions, not raw models they had to assemble themselves.
"We initially thought model capability was the product," Mei recalled in a May 2026 interview with 36Kr.
Tao Mei, Founder and CEO at HiDream.ai
This realization pushed HiDream.ai to build an agent or workflow layer between its foundation models and business users. The company experimented with different business models, including subscriptions, content services, and revenue sharing tied to gross merchandise value. E-commerce became a key testing ground: the company initially targeted conventional product imagery, but those assets remain unchanged for months, limiting recurring demand. It then shifted toward content-driven commerce and marketing, where brands may need thousands of short videos per month and AI becomes part of a continuous production workflow.
By the first quarter of 2026, HiDream.ai's products served more than 30 million professional users and more than 40,000 enterprise customers worldwide. The company's "1+1+3" strategy combines foundation models, an enterprise model-and-agent platform, and applications for marketing, film and television, and social media. The platform can also use third-party models when they better fit a specific task.
Steps to Understanding HiDream.ai's Competitive Advantage
- Cost Efficiency: HiDream.ai reduced video training costs to roughly one-fifth of the industry average through its dual-model approach, testing ideas on cheaper image models before scaling to video.
- Architectural Innovation: The company's Unified Transformer architecture processes multiple data types (text, images, video, spatial information) in a single framework, enabling interactive world models that preserve spatial consistency during navigation.
- Enterprise Focus: Rather than competing on raw model capability alone, HiDream.ai built workflow and agent layers that integrate models into business solutions, targeting recurring content demand in marketing and e-commerce.
- Legal Data Foundation: The company accumulated 200,000 hours of licensed video through partnerships, addressing copyright risks that plague competitors in the AI video space.
How Does HiDream.ai Compare to Google DeepMind?
Google DeepMind offers a broader contrast to HiDream.ai's focused approach. Google's visual-AI portfolio spans Nano Banana and Imagen for images, Veo for video, and Genie 3 for interactive world models, all supported by Google's existing consumer, developer, and cloud ecosystem. Google can leverage its massive distribution channels and existing customer relationships to deploy new models quickly.
HiDream.ai has taken a different route. Without comparable distribution channels, the company treated enterprise services as a core part of its strategy from the start. The challenge was determining how to serve businesses through software subscriptions, content services, or models tied more directly to customer results. That focus is reflected in its model-platform-application stack and its move from low-frequency product imagery toward recurring content demand.
In June 2026, HiDream-O1-Image-1.5 reached the top three on Artificial Analysis' text-to-image leaderboard, ahead of Google's Nano Banana 2 at the time. By September 2, it ranked number 16, demonstrating how quickly technical leads can narrow in the competitive AI landscape. The company's API costs about $80 per 1,000 images, positioning it as a cost-conscious option for developers and businesses.
HiDream.ai has also used open source to build developer attention. In May 2026, it released the 8-billion-parameter HiDream-O1-Image model and code under an MIT license, making the technology freely available to researchers and developers. Mei has said that DeepSeek, a Chinese AI company that released powerful models as open source, changed his thinking about open source strategy, even when the immediate commercial return is unclear.
What Does HiDream.ai's Funding Tell Us About AI Investment Trends?
In July 2026, HiDream.ai announced a 1.5 billion yuan Series C funding round, bringing its financing over the previous three months to more than 2.1 billion yuan (roughly $290 million to $410 million USD, depending on exchange rates). This capital influx reflects investor confidence in the company's approach.
Wang Bing, an investor at Oriental Fortune Capital, explained what attracted his firm to HiDream.ai: competitive foundation models built at lower cost, high research and development and capital efficiency, and fast translation into enterprise use cases. He also highlighted the company's legally licensed visual data as important in a field where copyright remains a significant risk.
Despite the funding success, the economics remain demanding. Mei told Xinhua that talent, data, and compute remain expensive and that HiDream.ai expects to reach monthly break-even in 2029. This timeline underscores that even well-funded AI startups face years of losses before achieving profitability.
HiDream.ai still carries Mei's research background. In a 2021 JD profile, he argued that scientists should have freedom to pursue work that was "the best" or "the first," while engineers focused on standardized products and services. By 2026, however, Mei was reconsidering HiDream.ai's low-profile approach, including the need to strengthen its brand narrative. "We come from a scientist-entrepreneur background and are used to keeping our heads down and doing the work," he said. This shift suggests that as HiDream.ai matures, it will need to invest more in communicating its achievements to global audiences, not just building better models.