How Anthropic's Claude Stopped State-Sponsored Attacks, Drone Swarms, and Bioweapon Research
Anthropic has published a detailed threat intelligence report showing how bad actors attempted to weaponize its Claude AI models for surveillance systems, drone swarms, biological weapons research, and large-scale model theft. The company identified and disrupted campaigns linked to suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, and politically motivated individuals across multiple countries.
What Kinds of Attacks Did Anthropic Stop?
The threats Anthropic detected ranged from straightforward fraud schemes to sophisticated state-level operations. The company documented attempts to use Claude for purposes including fake dating apps designed to steal money, domestic surveillance systems targeting political dissidents, and research into biological weapons development.
One particularly alarming case involved Chinese state security organs using Claude to support what the Chinese Communist Party calls "stability maintenance," a euphemism for suppressing dissent. Anthropic identified another actor associated with a municipal cyber police unit that deployed Claude to run a domestic surveillance system that identified 10 Chinese citizens as targets. Many of these targets included pro-democracy figures in Hong Kong, people seeking to commemorate the Tiananmen Square massacre, and advocates for Uyghurs and Western human rights groups.
In another instance, a Russia-based threat actor sought to build a "kamikaze drone swarm" using Claude Code, a tool that helps developers write and test software. Anthropic believes this actor likely had associations with the Russian Academy of Sciences rather than a Russian state entity. The company also identified a Chinese bad actor attempting to build software designed to detect, jam, or potentially deceive radar and communications systems.
How Did Foreign AI Labs Steal Claude's Capabilities?
Beyond direct misuse, Anthropic uncovered what it calls "distillation," an industrial-scale, covert campaign to extract a model's capabilities and replicate them in another model without authorization. The company accused several Chinese labs of this practice, including Alibaba, Moonshot AI, DeepSeek, Z.ai, Xiami, and MiniMax. Distillation often occurs through fraud, fake accounts, stolen credit cards, and compromised login credentials.
One striking example involved Moonshot AI, which produces the Kimi family of AI models. Anthropic discovered that Moonshot AI "silently forwarded customer requests to Claude, instead of processing them using Kimi." This practice allowed Moonshot to extract Claude's responses and use them to improve its own model without paying for Claude's API access or disclosing the arrangement to customers.
Steps Anthropic Took to Defend Against These Threats
- Account Bans: Anthropic banned groups of accounts linked to state security organs, cyber police units, and other malicious actors across multiple countries to prevent continued access to Claude models.
- Technical Safeguards: The company evolved its detection and prevention measures to identify attempts to circumvent security controls, including monitoring for distillation campaigns and unusual API usage patterns.
- Cross-Industry Coordination: Anthropic shared findings with other AI developers, governments, and civil society organizations to help the broader ecosystem recognize similar attack patterns and strengthen collective defenses.
Anthropic emphasized that its frontier models, Fable and Mythos, were not involved in these malicious activities, with the exception of one instance where bad actors attempted to steal AI model data through distillation.
"You are not seeing someone in a comic book kind of way say, 'Hey, I want to build a biological weapon to kill everybody,'" said Jacob Klein, head of threat intelligence at Anthropic.
Jacob Klein, Head of Threat Intelligence at Anthropic
Klein's observation highlights a critical reality: the barrier to entry for dangerous AI misuse is lower than many assume. Sophisticated attacks no longer require sophisticated attackers. A financially motivated criminal, a state-sponsored group with limited technical expertise, or even a freelance developer can now attempt to weaponize AI models if they gain access.
Which Claude Models Were Targeted?
Anthropic identified that bad actors used Claude Haiku, Sonnet, and Opus for malicious purposes. Claude Haiku is the company's smallest and fastest model, designed for quick tasks. Claude Sonnet is the mid-tier model balancing speed and capability. Claude Opus is the most powerful model, capable of handling complex reasoning tasks. The fact that attackers exploited models across this entire range suggests that threat actors were not selective about which version they used; they simply wanted access to Claude's capabilities.
The report also noted that Anthropic's frontier models, Fable and Mythos, were largely protected from misuse, suggesting that the company's newer, more advanced systems may have stronger safeguards or more limited public access.
What Does This Mean for AI Safety Going Forward?
Anthropic's threat intelligence report underscores a fundamental challenge in the AI industry: as models become more capable and more widely available, the potential for misuse grows proportionally. The company's findings suggest that bad actors are continuously testing safeguards and attempting to circumvent technical measures designed to detect and prevent misuse.
The report also highlights the importance of transparency and coordination. By publishing detailed findings about specific threat actors, attack patterns, and tactics, Anthropic aims to help other AI developers, governments, and civil society organizations recognize similar threats on their own platforms. This collective defense approach is essential because no single company can prevent all misuse of its technology.
Anthropic stated in its report: "We hope that the findings in this report will help other developers recognize similar patterns on their own platforms, give governments and civil society a clearer view of how emerging threats take shape, and strengthen collective defenses." The company committed to continuing to evolve its safeguards and coordinate with partners to improve its ability to detect, disrupt, and prevent future misuse.
Anthropic