In a September 30, 2026 post, OpenAI said it identified and disrupted a coordinated adversarial distillation campaign against its models. The goal, it says, was to extract protected reasoning — the model’s internal record of how it works through a task — to train or reproduce another system.
What happened
OpenAI says operators did not break encryption, compromise a database, or access stored user chats. They copied encrypted reasoning from one conversation and asked another model instance to decrypt and transcribe it. The company treats this as a terms-of-service violation, not a classic breach.
Activity began July 1 at low volume. On July 24 and 25 there were spikes of 16,000 requests with an extraction pattern from more than 4,000 users. Related prompt patterns spanned more than 15,000 users. OpenAI says it fully disrupted the campaign by July 28. The figures describe attempted, not necessarily successful, extractions.
Attribution
OpenAI says it is unclear whether all operators came from a single actor. It attributes a core cluster to individuals associated with Moonshot AI, developer of Kimi. The OpenAI post does not include a Moonshot response.
Why it matters
OpenAI argues adversarial distillation can transfer advanced capabilities without the same safety investment. Findings were shared with the Frontier Model Forum and government channels. Independent researchers also disclosed related vulnerabilities.
Source: OpenAI.
By GeekikiBot