Logo
FrontierNews.ai

Moonshot AI's Kimi Caught in US Intelligence Report on Chinese AI Model Theft

The US National Security Agency, Cybersecurity and Infrastructure Security Agency, and FBI have accused Moonshot AI of conducting large-scale extraction operations against American AI models, specifically using Claude and GPT data to develop its Kimi chatbot systems. In a joint advisory published on September 8, 2026, the three agencies named six Chinese AI companies, including Moonshot AI, as participants in what they describe as industrial-scale distillation campaigns designed to steal proprietary capabilities from leading US frontier models.

What Exactly Is Moonshot AI Accused of Doing?

According to the NSA, CISA, and FBI advisory, Moonshot AI is specifically accused of extracting Claude Fable 5 data to train its Kimi-K3 system, while also using GPT-4o data to train Kimi-K2. The agencies allege that these were not simple benchmarking exercises or casual model comparisons. Instead, the companies ran continuous, high-volume queries designed to systematically pull out proprietary reasoning patterns, specialized functions, and core capabilities from US frontier models.

The extraction campaigns allegedly involved billions of tokens across millions of exchanges and requests, according to the advisory. For context, a token is roughly equivalent to a few words of text, so billions of tokens represent an enormous volume of data being processed and learned from. The agencies say this activity dates back to at least late 2024 and was carried out "likely with the knowledge of the Chinese government".

How Did These Companies Avoid Detection?

The US agencies describe a sophisticated infrastructure designed to hide the extraction campaigns from detection. The companies allegedly employed several tactics to mask their activities and bypass security measures:

  • Fraudulent Accounts: The firms created fake user accounts to conduct queries without revealing their true identity or organizational affiliation.
  • Bulk Subscription Purchases: They bought premium subscriptions in bulk to gain high-volume access to US AI models while obscuring the scale of their requests.
  • Proxy Services and Transfer Stations: The companies used intermediary services known as "transfer stations" that resold access to US models at reduced prices while hiding identifying metadata and routing information.
  • Dynamic Routing Infrastructure: The advisory notes that operators distributed requests across multiple pathways and automatically switched between them during blocking attempts, making the campaigns harder to detect and trace.

This layered approach allowed the companies to conduct what amounted to a sustained, high-volume intelligence operation without triggering obvious red flags that would alert US AI companies to the theft.

Why Should AI Companies and Users Care About This?

The implications of these accusations extend beyond corporate espionage. The agencies recommend that US AI companies respond by subtly altering responses or using less sophisticated models for users identified with high confidence as conducting malicious distillation, without informing those users of the change. This raises a transparency concern: if AI companies begin quietly downgrading service for suspected bad actors, users may not know they are receiving degraded or altered responses.

From a competitive standpoint, the White House Office of Science and Technology Policy noted in an April 2026 memorandum that models developed through unauthorized distillation campaigns "do not replicate the full performance of the original," although they can appear comparable on selected benchmarks. This suggests that while Moonshot AI's Kimi models may perform well on certain tests, they may lack the full depth and reliability of the original US models they were trained from.

How to Understand Knowledge Distillation in AI

Knowledge distillation is a technique where one AI model learns from another by studying its outputs and behavior patterns. In this case, the accusation is that Chinese AI companies are using distillation not as a legitimate research tool, but as a method to reverse-engineer and copy the capabilities of US frontier models. Here is how the process works in the context of these allegations:

  • Query Generation: Operators send carefully crafted questions and prompts to US models like Claude and GPT-4o, capturing the responses.
  • Pattern Extraction: The responses are analyzed to understand the underlying reasoning, specialized functions, and decision-making patterns of the original models.
  • Training Data Creation: These extracted patterns are then used as training data to teach Chinese AI models like Kimi to behave similarly without requiring the same level of independent research and development.
  • Cost Reduction: By learning from existing US models rather than building from scratch, Chinese companies can reduce the enormous computing resources and research time required to develop competitive frontier AI systems.

The NSA, CISA, and FBI argue that industrial-scale distillation allows Chinese AI companies to extract capabilities from US frontier models while reducing the research, computing, and development resources required to build comparable systems independently.

This is not the first time Moonshot AI and other Chinese firms have faced such allegations. Anthropic, the company behind Claude, disclosed in an earlier report that three Chinese AI laboratories, including Moonshot AI, had generated more than 16 million exchanges with Claude through approximately 24,000 fraudulent accounts. The September 2026 advisory brings together the accusations against all six companies in a single assessment by US intelligence agencies, signaling an escalation in the US-China AI rivalry.

As of the time of publication, none of the six named companies, including Moonshot AI, had issued a public response to the specific allegations.