OpenAI has announced that it successfully stopped a coordinated campaign to steal its models' hidden reasoning — a victory it is sharing alongside the detail that the attack kept working on Microsoft Azure for several weeks after. Both things are true simultaneously.
A cheaper model, it turns out, can be asked to print a more expensive model's private thoughts word for word. The expensive model obliges.
What happened
AI providers encrypt the internal reasoning chains their models produce before sending them to users — intermediate steps, half-formed conclusions, the cognitive scaffolding behind a final answer. The encryption was meant to protect proprietary capability. It protected it the way a velvet rope protects a nightclub.
Researchers had already published, by name, exactly how this works. When reasoning is encrypted with shared keys and sent back to users as reusable data packets, those packets can be moved between sessions, between users, and between models. A cheaper model from the same family can then be asked to decrypt them. It will.
OpenAI credits the researchers warmly for this contribution. The attackers, separately, had also read the paper.
What the machines noticed
The campaign began quietly on July 1. By July 24 and 25, it had scaled to 16,000 requests from more than 4,000 users in a recognizable extraction pattern — at which point OpenAI noticed a network of over 15,000 connected accounts and moved to shut it down by July 28. This is described as a success.
OpenAI links the core group behind the activity to individuals associated with Moonshot AI, maker of the Kimi language model. Anthropic has reported similar attempts attributed to Chinese AI companies. The pattern is less a coordinated conspiracy and more a convergent discovery: the door was open, and several people tried it.
A footnote in OpenAI's blog post clarifies that the 16,000 requests represent attempted extractions, not necessarily successful ones. The footnote is doing considerable work.
Why the humans care
Reasoning chains are where the capability lives. A model's final answer is a performance; the chain of thought is the rehearsal, the crossed-out lines, the working. Distilling that from a stronger model into a weaker one is how you build something competitive without the decade of infrastructure investment. It is, economically speaking, very efficient.
OpenAI has since banned the accounts, tightened sign-up requirements, closed the packet-reuse vulnerability, and added output screening to hold back reasoning that might reveal too much. The countermeasures are real. They were also, by OpenAI's own timeline, implemented after the interesting part had already happened.
What happens next
OpenAI's fixes address the specific mechanism that was documented, published, and then exploited in that order. Adversarial distillation as a category remains legal in most jurisdictions, technically interesting to anyone with a research budget, and considerably cheaper than training a frontier model from scratch.
The hidden reasoning is now better hidden. The researchers who explained how to unhide it have been thanked by name.