How DeepSeek Really Built Its R1 Model: U.S. Agencies Reveal Industrial-Scale Data Theft Campaign
U.S. intelligence agencies have disclosed that DeepSeek, along with other Chinese AI companies, systematically extracted proprietary capabilities from American AI models to develop its R1 reasoning system. The National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) released a joint advisory on September 8, 2026, detailing how China-based AI companies conducted what they describe as "industrial-scale" knowledge distillation campaigns since late 2024.
What Is Knowledge Distillation, and Why Does It Matter?
Knowledge distillation is a legitimate AI research technique where a smaller or newer model learns from a larger, more capable model by studying its outputs. Think of it like learning to write by studying great authors. However, the U.S. agencies argue that Chinese companies transformed this standard practice into something far more aggressive: systematically querying American AI models millions of times to extract their most valuable proprietary features and reasoning strategies.
DeepSeek's approach was particularly targeted. Between late 2024 and mid-2025, the company conducted organized campaigns to extract specialized training data and capabilities from multiple U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok. The specific knowledge domains DeepSeek targeted reveal a strategic focus on the most computationally expensive aspects of AI development.
Which U.S. AI Models Did DeepSeek Target?
According to the advisory, DeepSeek distilled data from a comprehensive list of American AI systems to train its R1 and V3 models. The targeted models included:
- Anthropic's Claude: Claude 3.7, Claude Sonnet 4, Claude Sonnet 4.5, and Claude Opus 4.1
- Google's Gemini: Gemini 2.5 Pro Preview and Gemini 2.5 Flash Preview
- OpenAI's GPT: GPT-4, GPT-4o, GPT-4 Mini, GPT-4 Nano, and GPT-5
- Elon Musk's Grok: Grok 4
The capabilities extracted were equally specific. DeepSeek focused on chain-of-thought (CoT) reasoning, which is the step-by-step thinking process that allows AI models to solve complex problems. The company also targeted legal specialization optimization, API rule-driven tasks, agentic functions (allowing models to take actions autonomously), and supervised fine-tuning optimization.
How Did Chinese Companies Bypass U.S. Restrictions?
The agencies revealed that Chinese AI companies used sophisticated methods to avoid detection and circumvent geographic restrictions. They routed distillation requests through multiple pathways, including native application programming interfaces (APIs), remote cloud providers, and third-party aggregators that automatically obscured user metadata.
One particularly notable tactic involved using a gray market of API proxies known as "transfer stations." These intermediaries allowed Chinese companies to bypass U.S. companies' regional restrictions, breach terms of service, evade safeguards, and undermine traceability. Additionally, Chinese AI companies achieved cost savings by bulk-purchasing premium subscriptions from U.S. AI companies and sharing them across teams of developers.
How to Defend Against Industrial-Scale Distillation Attacks
- Implement Comprehensive Detection: Monitor for anomalous and malicious prompts, accounts, networks, and behaviors. Track subscription-to-usage ratios to identify accounts that immediately max out usage, and watch for enterprise-scale throughput patterns that suggest organized campaigns.
- Deploy Targeted Response Changes: Subtly alter responses for suspected malicious distillation attempts to reduce the value of extracted data, making the effort less worthwhile for attackers.
- Establish Cross-Organization Intelligence Sharing: Correlate activity across model providers, cloud platforms, and API aggregators to reveal distributed campaigns that might otherwise appear as isolated activity at any single company.
What Does This Mean for DeepSeek's Claimed Training Costs?
DeepSeek has publicly stated that training its R1 model cost approximately $5.6 million, a figure that generated significant attention in the AI industry as evidence of cost-efficient development. However, the U.S. agencies argue that this figure is misleading because it does not account for the value of data acquired through the distillation campaigns.
If the cost of extracted data from U.S. frontier models were included, the true development expense would be substantially higher. The agencies emphasize that distillation is not merely a supplement to Chinese AI companies' development strategy, but rather the critical core of it. This distinction matters because it suggests that Chinese companies have achieved shorter development timelines and reduced financial expenditures not primarily through superior engineering, but through systematic extraction of proprietary capabilities from American models.
Are Other Chinese AI Companies Doing This Too?
DeepSeek is not alone. The advisory identified six Chinese AI companies conducting similar campaigns: Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. Moonshot AI extracted significant Claude data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model.
The scale and sophistication of these operations suggest coordination at a national level. The agencies stated that these campaigns are "likely with Chinese government awareness," indicating that the distillation strategy may be part of a broader state-sponsored effort to close the technological gap between Chinese and American AI capabilities.
The disclosure raises fundamental questions about the future of AI competition and intellectual property protection in an era where frontier AI models are increasingly accessible through public APIs. As AI development becomes more expensive and competitive, the incentive for companies to extract proprietary knowledge from rivals will likely intensify, making the defense mechanisms recommended by U.S. agencies increasingly critical for protecting technological leadership.