OpenAI Disrupts Coordinated Campaign to Extract Protected AI Reasoning
OpenAI says it has identified and shut down a coordinated effort to extract the protected internal reasoning of its AI models — the step-by-step "thinking" process that normally stays hidden from end users.
What happened
According to OpenAI, a core cluster of the activity traces back to the first week of July 2026 and has been linked to individuals associated with Moonshot AI, a Beijing-based AI company — though OpenAI has not published technical evidence for that attribution, which it attributes to security considerations. Crucially, OpenAI says no encryption was broken and no database or stored conversation was directly accessed. Instead, the operators manipulated model interactions at scale so that protected reasoning was reproduced in a form visible to the requester, in violation of OpenAI's terms of service.
The activity started at low volume on July 1, then spiked on July 24–25 to roughly 16,000 suspicious requests from more than 4,000 accounts using a specific extraction pattern. A broader investigation subsequently turned up related "prompt-pattern" activity across more than 15,000 users before OpenAI says it fully contained the campaign on July 28.
OpenAI describes the technique as "adversarial distillation": systematically harvesting one model's outputs to train, reproduce, or improve a separate model without authorization. In response, the company banned the accounts involved, closed a pathway that let someone already in possession of another user's encrypted reasoning replay and recover it, and added new checks to detect and withhold streamed output that risked exposing reasoning.
Why it matters
The timing lines up with academic research published in August 2026 by a team from MATS Research, the ELLIS Institute Tübingen, and Synk, who found that encrypted reasoning traces can be fully interchangeable across different sessions, users, and models within the same provider's ecosystem. In their words, feeding an encrypted reasoning trace from a stronger model into a weaker, less-guarded model from the same provider can force the weaker model to decode and output the trace in plaintext — without ever directly jailbreaking the original model.
That has implications well beyond one vendor. The same mechanism could enable large-scale private data extraction, let attackers smuggle hidden prompt-injection payloads inside encrypted reasoning blocks, and surface sensitive content that a model's visible, user-facing output would otherwise have refused to show. OpenAI itself frames adversarial distillation as a safety and national-security concern: reasoning extracted this way can be used to train other models without the safeguards applied to the original model's outputs, accelerating the transfer of advanced capabilities — a risk that grows sharper as models take on more dual-use tasks.
This is also not an isolated incident for Moonshot AI. Last month, Anthropic separately alleged that Moonshot AI was quietly routing customer requests for its Kimi model through Claude and returning Claude's responses to users, while retaining a subset of those exchanges to help train its own chain-of-thought model — activity that has been tracked under the identifier GTG-16002.
What to do
- Inventory your LLM dependencies. Know which reasoning-capable models your products or workflows call, directly or through a third-party integration, and what each vendor's terms say about distillation and reasoning exposure.
- Treat "encrypted reasoning" as a feature, not a security boundary. Don't assume opaque or encrypted model output is immune to extraction or replay — architectural flaws in how providers handle these traces have already been demonstrated.
- Monitor for anomalous usage patterns. If you operate or resell access to a reasoning model, watch for coordinated, scaled request patterns consistent with extraction attempts rather than normal usage.
- Review vendor risk for AI-to-AI integrations. If a provider relays requests through a third model (as alleged in the Moonshot/Anthropic case), confirm what data retention and training use that entails.
- Stay alert to prompt-injection delivered via encrypted or hidden channels, not just plain-text prompts, when assessing LLM-integrated applications.
