OpenAI says Moonshot AI led campaign to extract hidden model reasoning

Coordinated effort in July targeted protected chain-of-thought traces; no data breach confirmed

By LineZotpaper
Published
Read Time2 min
Sources2 outlets
OpenAI has revealed that a coordinated campaign in July attempted to extract protected reasoning traces from its large language models, with the company saying the effort was linked to actors associated with China-based Moonshot AI. The campaign, which OpenAI described as consistent with adversarial distillation, was disrupted by July 28, and the company stated that no database was compromised and encryption was not broken.

In a company blog post, OpenAI detailed a series of events beginning July 1 with low-volume activity that escalated to high-volume spikes on July 24 and 25, during which 16,000 requests using an extraction pattern were observed from over 4,000 users. The company said it was not certain whether all operators were the same actor or whether the extraction attempts were successful.

The method involved encrypted reasoning blocks that clients send back with each request. The operators reportedly tried taking encrypted reasoning from one conversation and asking a model in a different session to decrypt it. OpenAI said encryption was not broken, but it closed a pathway that allowed replay of another user’s encrypted reasoning to recover its contents. Additional checks were added to detect streamed output that might expose reasoning, and protections were strengthened across users, workspaces, organizations, and model families. OpenAI also worked with third-party providers to disrupt accounts whose activity passed through their services.

Independent security researchers had previously reported related vulnerabilities through responsible disclosure, and OpenAI confirmed the attack paths were real. A research paper dated Aug. 10 described testing on OpenAI, Anthropic, and Google models, where a weaker model was prompted to write out encrypted reasoning from a frontier model in plain text. The researchers said that after the providers acknowledged their report, they were unable to launch the same attacks.

Anthropic, in a September 2026 report, stated that Moonshot relayed almost 300,000 customer requests using a proxy network of 5,380 fraudulent accounts within a single ten-day period. Anthropic also said Moonshot saved Claude’s reasoning signatures and used cross-session replay attacks to convert them back into full reasoning traces.

In July, Moonshot denied that its Kimi K3 model was created from a distillation of OpenAI’s models, and at the time OpenAI President Greg Brockman said it was "too early" to determine whether distillation had occurred. OpenAI said its next steps include ensuring partner-hosted deployments have the same protections as first-party tools.

§

Analysis

Why This Matters

  • The incident highlights ongoing attempts by competing AI labs to extract proprietary model reasoning, which could allow rivals to replicate advanced capabilities without incurring the cost of original development.
  • It underscores a growing security challenge for AI providers: protecting not just model weights but the internal reasoning traces that reveal how models arrive at answers.
  • The involvement of an established Chinese AI company could escalate tensions around intellectual property protection and technology transfer between the US and China.

Background

OpenAI and other leading AI labs have long faced the risk of model distillation, where outputs are used to train competing models. The company implemented encrypted reasoning blocks to hide its models’ chain-of-thought from clients. The recent campaign attempted to circumvent those protections. Moonshot AI is a Beijing-based company known for its Kimi line of models, and has previously denied allegations of unauthorized distillation. The incident follows a pattern of reported attempts by Chinese entities to access Western AI technologies.

Key Perspectives

[OpenAI]: Sees the campaign as a systematic attempt to steal protected intellectual property and has implemented technical fixes and worked with third-party providers to disrupt the activity. It shared findings with unspecified “relevant authorities.” [Moonshot AI]: Has not commented on the latest OpenAI post, but in July denied that its Kimi K3 model was built from distillation and described its models as independently developed. [Critics/Skeptics]: May question attribution given the difficulty of conclusively linking specific users to an organization. Some may argue that distillation of reasoning traces is a grey area in AI research, though OpenAI frames it as unauthorized exploitation.

What to Watch

  • Further technical disclosures from OpenAI, Anthropic, or other providers about similar campaigns.
  • Responses from Chinese authorities or Moonshot AI if formal allegations or regulatory actions follow.
  • Adoption of stronger model reasoning protections across the industry, especially in partner-hosted deployments.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.