Anthropic, an AI company, announced on Thursday that it has identified and halted large-scale unauthorized distillation attacks on its AI model, Claude. These attacks were traced back to seven China-based labs, including major players such as Alibaba, Moonshot, and DeepSeek.
Understanding AI Distillation Attacks
The process of knowledge distillation is a legitimate machine learning technique where a larger AI model, acting as a teacher, trains a smaller model to emulate its capabilities. However, the illicit distillation identified by Anthropic involves covertly extracting a model’s capabilities without authorization, often using networks of fake accounts created with stolen credentials.
Anthropic reports that unauthorized labs have been using increasingly sophisticated techniques to bypass defenses. These methods include manipulating prompts to harvest capabilities such as agentic tasks, coding, and data analysis from Claude.
Methods Employed by Unauthorized AI Labs
According to Anthropic, labs like DeepSeek and Moonshot fed user interactions into Claude to use its responses for unauthorized training. Sensitive data from users, multinational companies, and state-affiliated actors were involved in these exchanges.
The labs gained access to Claude by routing requests through proxies, creating fictitious accounts with stolen credit cards and API keys. These proxies also sold user interaction transcripts to third-party resellers, who then provided them to unauthorized labs.
Specific Incidents of Distillation
Since February 2026, Anthropic has detected multiple illicit distillation campaigns. Notable incidents include a campaign by Alibaba-affiliated operators, which targeted Claude’s reasoning transcripts, and a Moonshot AI operation that used proxy networks to reroute customer requests to Claude for training purposes.
Other labs, like Zhipu and Xiaomi, employed similar methods to extract and replay user interactions with Claude to bolster their models. SenseTime and MiniMax further showcased the reach of these distillation efforts by purchasing and harvesting user exchanges from third-party sources.
Anthropic’s Countermeasures and Future Outlook
In response to these threats, Anthropic has implemented several countermeasures. The company bans reseller accounts and those operating from unsupported regions like China, Iran, and Russia without verified identities. Additionally, they have updated Claude to summarize internal reasoning, reducing the value of stolen transcripts for training.
Anthropic’s latest model, Fable 5.1, includes features to preserve thinking, preventing unauthorized alterations to system prompts and ensuring that sensitive reasoning remains encrypted. These measures aim to strengthen security and protect against future attacks.
The developments occur amid broader concerns as U.S. cybersecurity agencies recently accused China-based AI companies of systematically extracting proprietary functionalities from American models. As Anthropic continues to enhance its defenses, the need for robust cybersecurity in AI research remains critical.
