How China's AI Labs Are Quietly Buying American Training Data to Catch Up
Chinese artificial intelligence companies have discovered a backdoor to American technological advantage: they're simply buying the specialized training data that powers frontier AI models, while simultaneously stealing proprietary capabilities through systematic extraction campaigns. Federal cybersecurity agencies and policy experts warn that this dual strategy is accelerating China's AI progress and exploiting a critical gap in U.S. export controls that focuses on computer chips but ignores data providers entirely.
The concern intensified after Moonshot AI released Kimi K3, a Chinese model that pulled within striking distance of leading U.S. frontier models. The White House began weighing plans to ban enterprise use of Chinese models, and Congress escalated probes into Chinese AI adoption by American companies. Yet policymakers have largely overlooked how Chinese labs are accessing the most sophisticated training data in the world through a handful of American firms.
What Is Training Data and Why Does It Matter?
Training data is no longer raw internet text scraped from websites and forums. The frontier has shifted dramatically. Today, when an AI lab wants to improve its model at a specific task, it hires specialized data providers to build custom training environments where the AI practices and learns. These aren't simple datasets; they're engineered ecosystems.
Consider how a data provider might train an AI to perform medical diagnosis or professional consulting work. They assemble subject matter experts, build simulated workplaces with realistic software tools and internal files, create task suites, and design verification systems to measure success. A recent evaluation suite built by Mercor, one of the leading data providers, assembled 256 professionals averaging nearly 13 years of experience, including former consultants from McKinsey and BCG, bankers from Morgan Stanley and Citigroup, and corporate lawyers from Disney. These experts spent five to ten days constructing simulated firms where AI agents could be trained on authentic professional work.
The market reveals just how valuable exclusive access is: environments sold exclusively command four to five times the price of those sold nonexclusively. This concentrated market means that more than 75 percent of the training data business belongs to just four U.S. companies. Switching vendors is prohibitively expensive because maintaining quality while scaling is the number one bottleneck in the industry.
How Are Chinese AI Companies Accessing American Training Data?
According to recent reporting, the top six Chinese AI labs spend over $500 million annually with U.S. data providers. SemiAnalysis reported in January that Surge AI sells training environments to Chinese labs including Moonshot and Z.ai, and that this access played a significant role in increasing capabilities for Kimi K2 Thinking and GLM-4.6.
The arrangement is straightforward: Chinese companies simply purchase access to the same high-quality training data that American AI labs use. Unlike semiconductor chips, which face strict export controls, training data has virtually no regulatory restrictions. U.S. data providers including Scale AI, Surge AI, and Mercor work directly with AI labs to build specialized training data, and they have no legal obligation to refuse Chinese customers.
This creates a paradox. The U.S. government has spent years restricting China's access to advanced computer chips through export controls on companies like Nvidia and Intel. Yet the data pipeline, which experts describe as equally or more strategic, remains completely unregulated. China itself has recognized the strategic importance: Beijing's commerce ministry has reportedly begun consulting Alibaba, ByteDance, and Zhipu about restricting the transfer of their key training data out of China.
What Is Industrial-Scale Distillation and Why Is It Alarming?
Beyond purchasing training data, Chinese AI companies are conducting what federal agencies call "industrial-scale knowledge distillation" campaigns. This is not the legitimate research technique of the same name; it is systematic theft of proprietary capabilities from U.S. frontier models.
The National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) released a joint cybersecurity advisory on September 8, 2026, detailing how Chinese AI companies including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI have extracted billions of tokens across millions of requests from U.S. frontier AI models since at least late 2024. These companies targeted variants of Claude, GPT, Gemini, and Grok to extract proprietary functionalities and capabilities.
Moonshot AI specifically conducted widespread distillation campaigns against U.S. frontier AI companies since at least mid-2025. The company extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model. The distillation targeted supervised fine-tuning optimization, reinforcement learning, software engineering, and mathematical capabilities.
Chinese companies route these distillation requests through multiple pathways to avoid detection. They use native application programming interfaces (APIs), remote cloud providers, third-party aggregators that automatically obscure user metadata, and a gray market of proxies known as "transfer stations" that bypass geographic restrictions and breach terms of service. They also achieve cost savings through bulk procurement of premium subscriptions shared across teams of developers.
How to Protect American AI Training Data: Policy Recommendations
- Extend Bulk Data Restrictions: The White House should extend the Department of Justice's existing bulk data restrictions, which already bar the sale of Americans' personal data to China, to cover AI training data as well. This regulatory framework already exists and can be adapted without creating new bureaucracy.
- Regulate Data Providers Directly: U.S. data providers including Scale AI, Surge AI, and Mercor should be restricted from selling to Chinese AI companies. High-quality training data is the scarcest ingredient in AI development, and these American firms should not subsidize Chinese AI progress on the backs of U.S. talent.
- Implement Detection and Mitigation: U.S. AI companies should detect anomalous and malicious prompts, accounts, networks, and behaviors. They should monitor subscription-to-usage ratios, immediate maximum usage from new accounts, and enterprise-scale throughput patterns to identify distillation campaigns.
- Deploy Targeted Response Changes: AI companies should subtly alter responses for suspected malicious distillation attempts to reduce the payoff to companies conducting industrial-scale campaigns, making the theft less valuable.
- Establish Cross-Organization Intelligence Sharing: Model providers, cloud platforms, and API aggregators should correlate activity to reveal distributed distillation campaigns that deliberately spread operations across multiple providers to avoid single-point detection.
Why Is China's Own Data Ecosystem Underdeveloped?
China's inability to build a comparable training data industry reveals why American data providers are so strategically important. A Z.ai cofounder complained in People's Daily that China's high-quality data is "fragmented and scattered." The Center for a New American Security found that China's data industry is immature enough that its labs resort to harvesting U.S. model outputs, which "spares Chinese developers' own limited compute for other uses," allowing China to be an "even faster follower".
Nathan Lambert, an AI researcher at the Allen Institute who visited most of the leading Chinese labs in spring 2026, observed that there was "almost no data industry" comparable to the United States'. Building high-quality training data requires both institutional knowledge and significant computing resources. Vendors now use frontier models to generate data and judge other models' attempts. Mercor's chief executive confirmed the company spends more on tokens, a measure of computing power, than on employee salaries.
What Are the Broader Implications for U.S. Technological Leadership?
The distillation campaigns have concrete consequences. DeepSeek claimed its R1 model cost only $5.6 million to train, a figure that seemed implausibly low. However, this calculation does not include the true cost of data acquired through extensive malicious distillation. DeepSeek distilled specialized training data and capabilities from Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, Claude Opus 4.1, Gemini 2.5 Pro Preview, Gemini 2.5 Flash Preview, GPT-4, GPT-4o, GPT-4 Mini, GPT-4 Nano, GPT-5, and Grok 4.
The specific capabilities extracted included legal specialization optimization, API rule-driven tasks, writing using chain-of-thought reasoning, agentic functions, question and answer optimization, coach and assistant capabilities, functional creation optimization, supervised fine-tuning optimization, and creative and occupational writing optimization.
Chinese AI companies that conduct industrial-scale distillation see significantly shorter AI development timelines and reduced financial expenditures in training frontier models. They deliberately distribute operations across multiple providers, platforms, and pathways to avoid single-point detection. This represents systematic extraction of proprietary functionalities and capabilities that threatens U.S. technological leadership.
The White House's AI Action Plan already recognizes high-quality training data as "a national strategic asset." The question now is whether policymakers will treat it with the same urgency they have applied to semiconductor export controls. Without action, American data providers will continue subsidizing Chinese AI progress, and the technological gap that the U.S. has worked to maintain will narrow further.