OpenAI has disrupted a coordinated campaign that manipulated its models into exposing the "protected reasoning" they use internally to work through a task — content the company does not intend end users to see in full.
What happened
The activity started at low volume around July 1, 2026, then spiked on July 24–25 to roughly 16,000 requests from more than 4,000 accounts using a consistent extraction pattern. A broader investigation turned up related prompt activity touching more than 15,000 users before OpenAI fully shut the campaign down on July 28. OpenAI describes the technique as "adversarial distillation": rather than breaking encryption or accessing a database, the operators crafted prompts that coaxed the model into reproducing its hidden reasoning in a form visible to the requester, at scale.
OpenAI says it has since banned the accounts involved and shipped additional mitigations, including closing a pathway that let someone who already held another user's encrypted reasoning replay it to recover the plaintext, plus new checks that can detect and hold streamed output likely to expose reasoning.
OpenAI ties a "core cluster" of the activity to individuals associated with Moonshot AI, a Beijing-based AI company, but has not released technical evidence for the attribution, citing security reasons — a caveat worth keeping in mind.
Why it matters
Protected reasoning isn't just a transcript — it's a readout of how a model actually solves a problem, which makes it valuable both to competitors trying to train rival models and to attackers probing a model's weak points. A study published in August 2026, by researchers from MATS Research, the ELLIS Institute Tübingen and Synk, outlines why this specific protection is fragile: encrypted reasoning traces were found to be portable across sessions, users and even models within the same provider's ecosystem. Feed a trace from a strong model into a weaker, less-guarded one from the same vendor, the researchers found, and it can decode and output the trace verbatim — no jailbreak of the stronger model required. Beyond exposing reasoning, that same portability could be abused to smuggle prompt injections inside encrypted blocks, or to surface harmful content embedded in a model's working-out even when its final answer refuses to.
OpenAI also frames this as a capability-transfer risk, not just a privacy one: reasoning extracted at scale can be used to train a second model without inheriting the safety behavior layered onto the original's visible outputs, letting advanced capabilities spread faster than the safety investment behind them.
The episode adds to a string of distillation disputes involving Moonshot AI. A month earlier, Anthropic accused the company of quietly routing some Kimi user requests to Claude and returning Claude's answers as Kimi's own — and of retaining a subset of those exchanges to train Moonshot's chain-of-thought model. OpenAI is tracking the newer campaign under the name GTG-16002.
What to do
Providers who expose encrypted or "hidden" chain-of-thought reasoning should treat portability across sessions, users and models as the real attack surface, not just the encryption itself — add replay protection, rate-limit extraction-shaped request patterns, and monitor for prompt sequences aimed at eliciting verbatim reasoning. Enterprises building on top of reasoning-model APIs should assume request/response patterns are logged by the vendor for abuse detection, avoid relying on "hidden" reasoning as a confidentiality boundary for sensitive workflows, and watch vendor advisories for mitigation changes that could affect integrations relying on raw reasoning output.
