TL;DR

In a disclosure published on Wednesday 30 September, OpenAI said it had found and shut down an organised attempt to extract the hidden reasoning its models produce while working on a task. The activity started on 1 July and was fully disrupted by 28 July. OpenAI pins a core cluster on people linked to Moonshot AI, the developer behind the Kimi models.

What OpenAI says happened

The company describes this as adversarial distillation: using another model’s outputs or reasoning, without permission, to train or improve a rival. Protected reasoning matters because it can reveal material deliberately left out of the final answer and makes a model’s skills easier to copy.

OpenAI is clear about what did not happen. Nobody broke its encryption, got into a database or reached stored user conversations. Instead, operators manipulated the models into reproducing protected reasoning in a form the requester could read. One method lifted encrypted reasoning out of a chat and asked the model, in a separate chat, to decode and write it out.

Volume was low at first, then spiked on 24 and 25 July with 16,000 requests from more than 4,000 users. OpenAI notes these were attempts, not necessarily successful extractions. Further digging linked prompt patterns across over 15,000 users. Outside security researchers separately reported related flaws through responsible disclosure, and OpenAI confirmed they were real.

The attribution is hedged. OpenAI says it is unclear whether one actor was behind everything it saw.

How it responded

OpenAI says it shut down or limited fraudulent accounts, tightened signup controls, and closed a route that let anyone holding another user’s encrypted reasoning replay it to recover the contents. Findings went to other labs via the Frontier Model Forum and to government information-sharing channels.

Why this matters now

It comes as Moonshot carries out an internal review of its Kimi models following a bioweapons jailbreak reported by Mindgard. In September Anthropic also described distillation cases involving Chinese labs. Taken together, leading labs are now publicly naming the rivals they blame, rather than treating extraction as an anonymous terms-of-service problem.

There is a two-month gap between disruption and disclosure, which OpenAI says was spent on mitigations and checks with partners.

Looking forward

OpenAI says models hosted by its partners must be protected just as well as its own services. For UK organisations reaching OpenAI models through cloud partners rather than directly, that is the practical question to raise with suppliers: which of these new controls apply to the version they are actually running.