Skip to main content
AI & Machine Learning

US Accuses China’s Moonshot AI of Illicit Model Distillation

Rows of server racks in a high-tech research facility hint at large-scale computing infrastructure.

“To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection,” Michael Kratsios wrote on X Wednesday.

Michael Kratsios and the White House allegation

Michael Kratsios, who leads the White House Office of Science and Technology Policy, publicly accused Beijing-based Moonshot AI of distilling Anthropic’s recently released Fable model to produce its own K3 model. Kratsios said the company built a “sophisticated internal platform” to run large-scale distillation against U.S. models and to evade detection, and that Moonshot had used GB300 servers — “either newly acquired or through Thailand” — to train its models. Kratsios framed the claim within a broader defense of competitive, legitimate AI development while denouncing covert industrial distillation: “Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”

Moonshot AI and the Kimi K3 model

Moonshot AI’s own website markets the Kimi K3 as “the first open 2.8 trillion parameter model,” touting lower token costs and “near-frontier performance.” The company says Kimi K3 “still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol,” but that it “demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.” A request for comment sent to Moonshot AI was not returned before the publication of the source article.

Anthropic, prior allegations, and the mechanics of distillation

Industry observers and security vendors describe distillation as a technique that can extract many of a larger model’s core capabilities. Piyush Sharma, CEO of Tuskira (an AI cybersecurity detection and response company), said distillation allows developers to replicate a model’s core abilities and pointed to an earlier complaint by Anthropic against Alibaba. According to Anthropic, that campaign used 25,000 fraudulent accounts to run 28.8 million interactions on Claude over six weeks — a volume Sharma said made “the goal was clearly replication.” Sharma warned that “when a model has learned to reason through software weaknesses, security gaps, and attack paths, copying its behavior also copies that analytical capability.”

Congressional scrutiny: Garbarino and Moolenaar’s joint inquiry

In April, Rep. Andrew Garbarino (R-N.Y.), chair of the House Homeland Security Committee, and Rep. John Moolenaar (R-Mich.), chair of the Select Committee on China, announced a joint investigation into the integration of Chinese AI models. The committees said the inquiry will focus on “examining a pattern of conduct by [Chinese]-based AI laboratories involving the large-scale theft of proprietary capabilities from American frontier AI systems through adversarial distillation,” the redistribution of those capabilities “as open-weight models available for global download,” and the “incorporation of PRC-origin models into products used daily by hundreds of thousands of American developers and engineers.” Frontier AI companies in the U.S. have pressed policymakers to make copying or duplicating advanced commercial models more difficult, calling such activity a form of intellectual property theft.

What this means for technologists, policymakers, and enterprises

  • Technologists and security teams: will be watching detection and mitigation techniques for large-scale distillation, and monitoring model access and training logs for high-volume or adversarial activity — concerns echoed by Tuskira’s CEO about replication through high-volume interactions.
  • Policymakers and investigators: have an active avenue of inquiry through the Garbarino–Moolenaar joint investigation and face a policy tradeoff between protecting proprietary capabilities and sustaining open innovation, a tension Kratsios acknowledged in his statement.
  • Enterprises and procurement leaders: must weigh claims about the provenance of open-weight models and the committees’ stated focus on PRC-origin models’ incorporation into widely used developer tools, even as some industry participants argue that many models were trained on broadly shared internet content.

The allegation places three specific facts at the center of the dispute: the White House claim that Moonshot distilled Anthropic’s Fable to build K3, Moonshot’s public positioning of Kimi K3 as an “open 2.8 trillion parameter model” with near-frontier performance, and the congressional investigation that will probe large-scale adversarial distillation and redistribution of open-weight models. Key evidentiary details remain undisclosed in the public record: Kratsios did not provide specifics on how the U.S. government determined that K3 was distilled from Anthropic, and Moonshot did not respond to requests for comment prior to publication. Those gaps — alongside Anthropic’s earlier allegations against Alibaba and the technical warnings from AI security vendors — ensure the dispute will be resolved in public inquiries, private forensic work, and contested industry debate.

Read the original CyberScoop report