"The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations."
How OpenAI says the attack worked
OpenAI described a coordinated campaign that sought to "distill and extract reasoning capabilities" from its models by copying encrypted reasoning data from one conversation and then prompting the model in a separate conversation to decrypt that content and transcribe it into plain text. Company engineers characterized the method as "novel" and said the attackers manipulated model interactions to reproduce protected reasoning in forms visible to the requester, in a way that violated OpenAI’s terms of service. OpenAI emphasized that operators did not break encryption, compromise a database, or gain direct access to stored user conversations, and framed the incident as an abuse of model interactions rather than a conventional data breach.
The timeline: July spike and disruption
OpenAI reported it first saw low-level activity on July 1 that gradually increased through the month. Activity spiked on July 24 and 25, dates when OpenAI observed 16,000 prompts from 4,000 users that matched a similar "relevant extraction pattern." By July 28, the company said the number of suspicious users had risen to 15,000, at which point OpenAI "fully disrupted" the operation. Outside researchers also reported a similar vulnerability to OpenAI in August.

Audit-ready is a season. It shouldn't be.
Evidence in spreadsheets, controls drifting between audits, frameworks multiplying on flat headcount. Nubivance runs continuous compliance on Rapid7 Cyber GRC - SOC 2, HIPAA, ISO 27001, PCI, CMMC.
End the scrambleAttribution to Moonshot AI and limits of technical citation
OpenAI said individuals working on behalf of Moonshot AI, a China-based rival, were behind a "core cluster" of the activity. The blog post did not, however, cite technical evidence or reasoning supporting that attribution. CyberScoop reported the attribution and noted OpenAI’s post was unsigned. OpenAI told CyberScoop it was not sharing additional information "for security reasons," and CyberScoop reached out to Moonshot AI for comment.
Mitigations: fixes, account bans, and information sharing
OpenAI described a multi-pronged response. The company said it banned the offending accounts, improved signup and infrastructure controls, and expanded network monitoring. It also fixed a bug that specifically allowed users to take encrypted data from one conversation and decrypt it in another. Beyond internal changes, OpenAI said it has shared information about the incident with external groups including the Frontier Model Forum, and warned that the same vulnerability exists in other AI models.
What this means for Moonshot AI, model operators, and enterprise security teams
- Moonshot AI and similar model operators: the company named in OpenAI’s post — Moonshot AI — and other open-source Chinese models such as Kimi were highlighted as part of the competitive landscape. OpenAI’s public attribution to individuals working on behalf of Moonshot AI places scrutiny on those actors, while OpenAI declined to publish additional technical detail.
- Model operators and platform security teams: OpenAI’s fixes — banning accounts, tightening signup controls, patching the cross-conversation decryption bug, and expanding network monitoring — are concrete steps security teams can expect platforms to adopt after similar incidents. OpenAI also reported sharing details with the Frontier Model Forum, indicating a channel for cross-platform coordination on such vulnerabilities.
- Enterprises and procurement leaders using large models: OpenAI warned the same vulnerability exists in other models. Enterprises that depend on externally hosted models may need to evaluate vendor controls around conversation isolation, account creation and bulk-account acquisition risks, and incident disclosure policies.
Cybersecurity researchers and vendors also weighed in on the broader pattern OpenAI described: according to the reporting, cybersecurity experts at Google and other firms say Chinese companies sometimes obtain thousands of individual accounts through black or gray markets and then flood models with millions of prompts and data requests to copy capabilities and training data. American AI companies and the U.S. government have accused firms like Moonshot AI of conducting "systematic" distillation attacks on recent models — allegations OpenAI reiterated in its public post while declining to make detailed technical evidence public.
OpenAI’s account sketches a hybrid threat: not a classical data breach, but a scalable manipulation of model interactions that reproduces protected reasoning outputs. The company says it disrupted a spike in activity and patched a cross-conversation decryption bug, and it has signaled the issue carries across multiple models. What remains publicly unresolved is the technical linkage supporting attribution to a single rival and whether independent researchers will publish corroborating analysis. The next clear step will be whether OpenAI, independent researchers, or affected platforms produce verifiable technical evidence that clarifies how the bypass worked and how widespread it may be.




