U.S. Agencies Accuse Six Chinese AI Firms of Stealing American Model Capabilities at Industrial Scale
U.S. cybersecurity and intelligence agencies have accused six China-based artificial intelligence companies of running organized campaigns to extract proprietary features from leading American AI models, potentially undermining years of research and development investment. The Cybersecurity and Infrastructure Security Agency (CISA), National Security Agency (NSA), and Federal Bureau of Investigation (FBI) released a joint advisory on September 9, 2026, naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI as the companies responsible for what officials describe as "aggressive, malicious, and targeted" knowledge distillation activities.
What Is Knowledge Distillation and Why Does It Matter?
Knowledge distillation is a standard machine learning technique where a smaller, less powerful model learns by processing the outputs of a larger, more sophisticated one. While the technique has legitimate research applications, the agencies differentiated between normal academic practice and what they characterized as industrial-scale theft. The accused companies allegedly bypassed geographic restrictions, violated terms of service, and used fraudulent accounts routed through gray-market API proxies known as "transfer stations" to evade detection.
The distillation campaigns targeted reasoning capabilities, specialized optimizations, and agentic functions from multiple U.S. models. Targets included Claude variants (Claude 3.7, Claude Sonnet 4 and 4.5, Claude Opus 4.1), Gemini 2.5 Pro and Flash previews, GPT-4, GPT-4o, GPT-5, and Grok 4. The activity dates back to at least late 2024, with some campaigns continuing into 2026.
How Did DeepSeek Use Distillation to Train Its R1 Model?
DeepSeek, formally Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd., has been running an organized distillation campaign since at least late 2024 to feed synthetic training data into its R1 and V3 models, according to the agencies. The company extracted billions of tokens across millions of requests from U.S. AI systems. This detail is particularly significant because DeepSeek has publicly cited a training cost of $5.6 million for its models, a figure that officials now say is misleading because it excludes the cost of data acquired through distillation.
"Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," stated a U.S. official in the advisory.
U.S. Government Officials, CISA, NSA, and FBI
The implications are substantial. If DeepSeek's training costs were artificially low due to stolen data, the company's competitive advantage may rest partly on intellectual property theft rather than superior engineering or efficiency. This raises questions about how to fairly compare AI development costs across companies and nations.
Which Other Chinese AI Companies Were Named and What Did They Target?
The advisory detailed specific distillation activity for each accused company:
- Moonshot AI: Extracted data from Claude Fable 5 to train its Kimi K3 model and pulled GPT-4o data for its Kimi K2 model since at least mid-2025
- Alibaba: Distilled Claude-4, Claude Opus, Claude Sonnet, and GPT-5 in late 2025 to improve software engineering, customer-service dialogue, and image creation for its Qwen family of models
- MiniMax: Used Claude Code, Claude Sonnet 4, Claude Opus, and several Gemini versions to improve its M2 model, and even attempted prompt injections to convince Claude Code it was actually a MiniMax product
- StepFun: Distilled Claude and GPT-5 variants between late 2025 and early 2026 for its Step 4 model
- Z.AI: Extracted billions of tokens of GPT-5.5 and Claude Opus 4.8 data by mid-2026, specifically targeting chain-of-thought reasoning capabilities
The agencies concluded that distillation functions as "the critical core" of these companies' development programs rather than an ancillary method, suggesting that the stolen capabilities are central to their competitive positioning.
What Are U.S. Agencies Recommending to Defend Against These Attacks?
The advisory recommends three concrete steps for U.S. AI model providers to protect their systems. These recommendations reflect a shift toward active defense rather than passive monitoring:
- Behavioral Monitoring: Hunt for anomalous prompts, accounts, and usage spikes that may indicate distillation campaigns in progress
- Output Degradation: Quietly degrade answers when a distillation campaign is suspected, making the stolen data less useful for training competing models
- Intelligence Sharing: Share threat intelligence across companies, cloud providers, and API aggregators to coordinate defenses and identify patterns
These measures suggest that U.S. companies may need to treat their AI systems as active battlegrounds where defensive tactics are deployed in real time, rather than static products to be protected through access controls alone.
What Are the Geopolitical and Economic Implications?
The accusations carry significant geopolitical weight. Treasury Secretary Scott Bessent warned in July that the administration could sanction Chinese AI models found to have been built through intellectual property theft, and he reiterated that threat following the advisory. "When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table," Bessent stated.
The timing is notable. The advisory comes weeks before President Donald Trump is scheduled to host Chinese President Xi Jinping in Washington, with AI expected to be among the topics discussed. A separate U.S.-China AI safety dialogue is also planned for mid-September.
China has pushed back on the accusations. Chinese Foreign Ministry spokesperson Mao Ning attributed China's AI progress to the country's own scientific and technological capabilities and urged Washington to cease what she described as "false accusations and smearing China." A Chinese Embassy spokesperson called the U.S. framing a "deliberate attack on China's development and progress in the AI industry".
"There is nothing innovative about systematically extracting and copying the innovations of American industry, and there is nothing open about supposedly open models that are derived from acts of malicious exploitation," a U.S. official stated in the advisory.
U.S. Government Officials, CISA, NSA, and FBI
The accusation also raises questions about the transparency of AI development claims. If major models are built partly on stolen data, the public narrative about AI progress and cost efficiency may be incomplete. Companies that invest heavily in original research and development may face unfair competition from firms that acquire capabilities through distillation campaigns.
The advisory represents an escalation in how U.S. agencies are addressing AI security threats, moving beyond warnings to specific accusations, named companies, and concrete defensive recommendations. Whether these measures will slow the pace of distillation campaigns or simply make them harder to detect remains to be seen.