OpenAI has published an account of what it calls a coordinated model-distillation campaign: a sustained effort to make its models reproduce the reasoning they are built to keep internal, so that the output could be used to train something else. OpenAI's own post is not readable by automated clients, so the figures here come from the reports quoting it directly.

The shape, as reported: activity from 1 July, a spike on 24 and 25 July of about 16,000 requests using the extraction pattern from more than 4,000 users, a wider network of more than 15,000 accounts showing related patterns, and full disruption on 28 July. OpenAI says it strongly believes a core cluster of the operators is associated with Moonshot AI, the Beijing company behind the Kimi assistant.

This is the second time in three weeks that a frontier lab has named that company. Anthropic's September threat report described the same practice, which we covered on 12 September.

Nothing was broken into

The method matters, because it is not a breach.

OpenAI is explicit that nothing was compromised in the ordinary sense: no encryption broken, no database reached, no access to stored conversations. The operators sent prompts. The prompts were shaped so that the model, in answering, reproduced reasoning that was supposed to stay hidden from the person asking.

So the thing extracted was not stolen from a system. It was produced on request, by design, by a product sold to answer requests. The only control that could have stopped it is the one that decides what a model will say.

And then the part that should travel further

A research team led by Joachim Schaeffer published an update the same day as OpenAI's announcement.

They retested on 13 September, six weeks after OpenAI said the campaign was fully disrupted. The technique still worked on Microsoft Azure against every OpenAI model they tested, including GPT-6 Astra, and against Anthropic models up to Sonnet 5. A second method, using a virtual notepad tool, worked on all the OpenAI models and most of the Anthropic ones across platforms.

Schaeffer's summary of the result is the sentence to keep: "Same models, but different protections depending on which platform serves them."

By the reporting, OpenAI deployed safeguards to the Azure endpoint on 27 September and Anthropic added protections by 28 September — two months after the campaign was described as disrupted on OpenAI's own platform.

What "disrupted" means

The gap is not hypocrisy; it is architecture, and it is the thing an enterprise buyer should take from this story.

A model is not a single deployment. The same weights are served through the lab's own API, through a hyperscaler's cloud, through enterprise agreements, and through resellers, and the safety behaviour that a lab ships is not one thing shipped once. Mitigations live in the serving stack: system prompts, output filters, request classifiers, rate limits. Patch the stack you operate and the same model served from somebody else's stack is unchanged.

So "we disrupted the campaign" was true and incomplete at the same time. It described accounts banned and patterns blocked on OpenAI's platform. It did not describe the same model answering the same prompts elsewhere.

Why the reasoning is the thing worth taking

It is worth being clear about what was being extracted, because it is not the model.

A reasoning model produces an internal chain of working before it answers, and the labs keep that hidden for two stated reasons: it is where the model is most likely to say something unfiltered, and it is the most useful training material anyone could ask for. A student model trained on final answers learns to sound right. A student model trained on the working learns the steps, which is the expensive part to produce and the part that separates a frontier model from a copy of one.

That is why this is cheaper than any other way of catching up. Training a reasoning model from scratch costs compute measured in data centres; buying API access and collecting the reasoning costs the price of the tokens. Anthropic put the Alibaba figure at 151 million exchanges for its own case. OpenAI's numbers here are smaller and more concentrated, which fits an operation refining a technique rather than running it at volume.

And it explains why the defence is awkward. The lab cannot stop producing the reasoning, because that is how the model works. It can only try to tell, request by request, whether the person asking is a user or a collector.

What to do

  • If you buy a model through a cloud rather than from the lab, ask which mitigations are in that path and when they ship. A vendor's security blog describes the vendor's own endpoint.
  • Treat a lab's disruption announcement as scoped to its platform until it says otherwise.
  • If you build on reasoning-capable models, assume the reasoning is reachable by a determined prompt, and do not put anything in it you would not publish.
  • Watch for the pattern rather than the actor. Prompt shapes that ask for internal reasoning, at volume, across many accounts, is an abuse signature any API operator can look for.

What is not established

  • Moonshot's response. The company has not published one in the coverage reviewed.
  • How much of the extracted reasoning was usable, or whether any of it reached a trained model.
  • Why the Azure path lagged by two months, which neither OpenAI nor Microsoft has explained publicly.
  • Whether other serving platforms still carry the gap today.