"The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations," OpenAI said.
OpenAI's account: discovery, scale, and disruption
OpenAI reported that it identified and disrupted a coordinated campaign aimed at illicitly extracting "protected reasoning" from its models. The company traced a "core cluster of the activity" to the first week of July 2026 and said the campaign began on July 1, 2026, initially at low volume before spiking on July 24–25 to some 16,000 attempted requests using a relevant extraction pattern from over 4,000 users. Upon further investigation, OpenAI identified related "prompt-pattern activity" across more than 15,000 users and said the campaign was fully disrupted on July 28, 2026.
Technique named: adversarial distillation and encrypted reasoning traces
OpenAI described the operation as adversarial distillation — the systematic, unauthorized use of one model's outputs to help train, reproduce, or improve another model. The company said operators "manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service."
A separate August 2026 study by researchers from MATS Research, ELLIS Institute Tübingen, and Synk identified an architectural vulnerability in encrypted reasoning traces that made them "fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem." The study warned that an attacker could exploit that compatibility to force a less-safeguarded model to decode and output a trace verbatim: "By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly."

This site is the portfolio.
OSINTSights runs on Cloudflare Workers, D1, R2, and Vectorize, with an AI pipeline on Hetzner ARM. Nubivance designed, built, and operates it. We do the same for clients.
See what we buildAttribution to Moonshot AI and prior allegations
OpenAI attributed a "core cluster" of the activity to individuals associated with Moonshot AI, a Beijing-based company, but did not cite technical evidence to support that assessment. The company emphasized that attackers did not break encryption or access stored conversations directly; rather, they manipulated model outputs to reveal protected reasoning.
This is not Moonshot AI's first public entanglement with distillation claims: last month, rival Anthropic accused Moonshot AI of relaying customer requests to Claude instead of processing them using Kimi, then displaying Claude's responses to users and retaining a subset of exchanges to train its chain-of-thought model. The activity has been tracked under the moniker GTG-16002.
OpenAI's mitigations and operational changes
OpenAI said it deployed additional mitigations, banned the fraudulent accounts it identified, and "closed a 'pathway' that made it possible for some who already possessed another user's encrypted reasoning to replay it and recover its contents." The company added checks to detect and hold streamed output that might expose reasoning, steps it described as part of its response to the campaign.
What this means for technologists, policymakers, and affected enterprises
- Technologists and security teams: The study's finding that encrypted reasoning traces can be compatible across sessions and models focuses attention on cross-model attack surfaces. Teams responsible for model deployments will need to monitor for prompt-pattern activity and replay techniques that attempt to use one model’s outputs to provoke another model into revealing hidden state.
- Policymakers and regulators: OpenAI framed the issue as a safety and national security concern, saying extracted reasoning "could be used to train another model without preserving the safeguards applied to the original model's user-facing outputs" and that "at scale, distillation can also accelerate the transfer of advanced capabilities." Regulators may be presented with technical questions about how to evaluate and certify protections that guard internal reasoning traces.
- Affected enterprises and procurement leaders: The incident highlights supply-chain and provenance questions for organizations that rely on third-party models. Claims that one provider's model outputs can be replayed into another model to recover internal traces makes contractual assurances about data handling and model isolation more salient.
OpenAI's public description leaves two facts central to the story: the company says it prevented a large-scale, coordinated extraction campaign that relied on manipulating model interactions rather than breaking cryptography, and an independent research group documented an architectural weakness that makes such manipulations feasible across sessions and models. Whether the architectural compatibility the researchers describe can be fully eliminated without degrading model functionality is an outstanding technical challenge, and the answers will determine how tightly providers must control model outputs, streaming behavior, and inter-model interoperability going forward.




