Skip to main content
Emerging Threats

OpenAI Disrupts Chinese AI Firm's Model Distillation Campaign

Rows of computer servers and networking equipment in a brightly-lit server room with a few workstation areas.

“The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations,” OpenAI wrote in a Wednesday blog describing an adversarial campaign it says reproduced its models’ protected reasoning at scale.

OpenAI’s account of the July distillation campaign

OpenAI says it detected a coordinated "distillation attack" that began on July 1 and ran through nearly all of July, culminating in a disruption on July 28. According to the company’s blog, the activity initially moved slowly but showed pronounced spikes on July 24 and 25, when the firm saw 16,000 requests matching a relevant extraction pattern originating from more than 4,000 users. Broader "prompt-pattern activity" linked to the campaign appeared across over 15,000 users.

OpenAI emphasized that the operators did not breach encryption, databases, or stored conversations. Instead, the company says actors "manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service."

Moonshot AI, Kimi, and the attribution OpenAI reported

While OpenAI said it could not definitively tie every operator involved to a single rival, the company identified a "core cluster" it attributes to Moonshot AI, the developer of a model called Kimi. The Register contacted Moonshot AI for comment and did not receive an immediate response, and OpenAI did not specify which of its models were targeted when asked.

The allegation fits within a string of similar claims: the article notes that federal authorities and major U.S. AI firms, including Google and Anthropic, have accused Chinese rivals — and specifically Moonshot AI — of using distillation to reproduce capabilities from American models. Separately, the article reports that Michael Kratsios, the U.S. President’s Assistant for Science and Technology, accused Moonshot AI in late July of creating its Kimi K3 model by distilling Anthropic’s Fable.

Model distillation, the risk OpenAI describes

Model distillation is described in the article as a technique that can use one model’s outputs to train another. In adversarial cases, attackers send bulk queries designed to recreate the larger model’s internal reasoning and capabilities in outputs they can collect. OpenAI framed this behavior as a safety and national-security concern: extracted reasoning, it said, could allow another model to be trained "without preserving the safeguards applied to the original model’s user-facing outputs," accelerating capability transfer without the same investment in safety, particularly in dual-use domains.

The Register also notes a parallel defensive move in the sector: Anthropic’s Claude Opus 5.5 model, released a week ago, includes a distillation-defense called "preserved thinking" that Anthropic introduced with Fable 5.1.

OpenAI’s mitigations and information sharing

According to the blog, OpenAI took several technical and operational steps to disrupt the campaign: it banned accounts tied to model-copying, tightened signup and infrastructure controls, expanded monitoring, closed "a pathway that allowed someone who already possessed another user's encrypted reasoning to replay it and recover its contents," and worked with service providers to prevent the activity from moving to third-party services.

OpenAI also shared the details of its investigation with other AI firms through the Frontier Model Forum and with government information-sharing programs, the company said.

What this means for technologists, policymakers, and competitors

  • Technologists and security teams: Expect continued focus on detecting bulk, patterned queries and on technical controls that prevent replay and extraction of encrypted or protected reasoning; the company reported specific spikes (16,000 requests on July 24–25) and an overall footprint exceeding 15,000 users tied to prompt-pattern activity.
  • Policymakers and regulators: The incident reiterates the concerns raised by U.S. agencies and officials about cross-border capability transfer; OpenAI labelled these risks as safety and national-security issues and has shared investigative details with government programs.
  • Competitors and model developers: Firms named in the article — Google, Anthropic, and Moonshot AI — are already embedded in the dispute: accusations and defensive responses (such as Anthropic’s "preserved thinking") signal an active technical and reputational contest over how models are built and protected.

OpenAI’s post lays out a concrete incident — dates, request volumes, mitigation steps, and a partial attribution to Moonshot AI’s Kimi — but it leaves at least one central fact publicly unresolved: which specific OpenAI models were targeted. That omission frames the next question for observers who must weigh the technical, safety, and policy implications of adversarial distillation across a rapidly evolving landscape.

Source: The Register — Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody else