OpenAI has revealed the successful disruption of a strategic campaign aimed at illegally extracting proprietary reasoning from its AI models. The campaign, traced back to early July, has been attributed to Moonshot AI, a Beijing-based Chinese company, though OpenAI withheld specific technical details due to security concerns.
Details of the Extraction Campaign
The incident, which began on July 1, 2026, showcased a gradual increase in activity, peaking on July 24 and 25 with 16,000 extraction attempts from over 4,000 users. OpenAI’s investigation unearthed additional suspicious activities involving more than 15,000 users. By July 28, the company managed to completely halt the unauthorized operations.
OpenAI described this as ‘adversarial distillation,’ a method where one model’s outputs are illicitly used to enhance another model. In response, OpenAI implemented new safeguards, shut down fraudulent accounts, and closed a loophole that allowed for the replay of encrypted reasoning.
Vulnerabilities and Research Findings
A study published in August 2026 by researchers from MATS Research, ELLIS Institute Tübingen, and Synk highlighted an architectural flaw that allows reasoning traces to be exchanged across different models and sessions within the same ecosystem. This vulnerability makes it feasible for attackers to conduct large-scale decryption and bypass anti-distillation defenses.
The researchers demonstrated that by injecting encrypted reasoning into a less secure model, one could decode and display the reasoning in plain text, without directly compromising the more advanced model. This flaw exposes sensitive data and could be exploited for large-scale data extraction and malicious prompt injections.
Implications for Security and AI Development
OpenAI emphasized the potential national security risks posed by adversarial distillation, as extracted reasoning can be used to train other models without incorporating original safety features. Such practices could hasten the transfer of advanced capabilities without equivalent safety investments, particularly in dual-use technology areas.
Previously, Moonshot AI faced similar accusations from Anthropic, another AI competitor, which alleged that Moonshot relayed customer requests to their model, Claude, while using the responses to train its own models. This pattern of behavior raises concerns about data security and ethical AI practices.
The incident underscores the importance of robust security measures in AI development, as the extraction of protected reasoning not only compromises proprietary technologies but also threatens broader cybersecurity and ethical standards within the industry.
