← News

OpenAI official announcement graphic titled Disrupting a coordinated model-distillation campaign

30 Sep 2026

OpenAI

OpenAI says it disrupted a coordinated model-distillation campaign tied to Moonshot

OpenAI said Wednesday it identified and disrupted a coordinated campaign to extract protected reasoning from its models — activity consistent with adversarial distillation — and attributed a core cluster of the activity to individuals associated with Moonshot AI.

Labs are racing to ship reasoning that stays hidden from the answer a person sees. Steal that chain, and someone can copy the capability without copying the safety stack that was supposed to stay on the original model. OpenAI’s same-day write-up treats that kind of distillation as a shared security problem, names a cluster tied to people associated with Moonshot, and says it has shared the findings with other labs and with government. That is a clear look at how competitive model copying actually shows up in the logs.

On Wednesday, 30 September 2026, OpenAI published that it identified and disrupted a coordinated campaign designed to extract protected reasoning from its models. The page is “Disrupting a coordinated model-distillation campaign.” It prints September 30, 2026. It does not print an hour. The earliest activity OpenAI says it observed was in the first week of July. Protected reasoning, as the page defines it, is the model’s internal record for working through a task. Pulling that record out can show information the final answer was supposed to withhold, and it can help someone else reproduce what the model can do. Those lines are OpenAI’s.

OpenAI says the activity is consistent with adversarial distillation. Distillation, in this use, means taking what one model produces and using it to train or improve another. Adversarial means the use was not authorized. The page’s definition is the systematic, unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Consistent with is the company’s wording for the activity it saw. The post does not say OpenAI recovered a finished copy of another model, and it does not print a dollar loss.

What the operators did not do, in OpenAI’s words. They did not break OpenAI’s encryption. They did not compromise a database. They did not gain direct access to stored user conversations. Encryption, here, is the lock that leaves data scrambled without the key. Instead, OpenAI says, they manipulated model interactions so protected reasoning could be reproduced in forms the requester could see, in a coordinated and scaled way that violated the terms of service. A term of service is the rule a person agrees to when they use the product. OpenAI says this kind of manipulation is not a weakness unique to its models, and that it shared information about it with industry partners through the Frontier Model Forum so the group could strengthen a common defense. The Frontier Model Forum is a group of companies that build the most capable models. Those lines are OpenAI’s.

One method the page describes. Operators tried to extract protected reasoning in new ways, including by copying encrypted reasoning from one conversation and asking a model in a different conversation to decrypt and transcribe the hidden reasoning. Decrypt means turn the locked text back into words. Transcribe means write those words out. The lock was not described as broken. The path OpenAI describes is a model being asked to read the locked text and write it out. Independent security researchers also brought related weaknesses to OpenAI through responsible disclosure, the practice of telling the company before telling the public. Those weaknesses crossed models, and they involved conversation compaction, which is squeezing a long chat into a shorter record the model can still use. OpenAI says it checked those findings, confirmed the paths were real, and that the researchers’ work helped it see the broader class of attack and speed up its fixes. The post does not name the researchers. Those lines are OpenAI’s.

The counts, and what they mean. OpenAI says the activity began on July 1, at a low volume at first. On July 24 and July 25 it saw high-volume spikes: 16,000 requests that used a relevant extraction pattern, from over 4,000 users. Sixteen thousand is the number of requests. Over 4,000 is the number of users on those two days, so the requests outnumber the accounts. Further investigation found related prompt-pattern activity across a cluster of more than 15,000 users. That cluster is wider than the 4,000 users on the spike days. A prompt is the text sent to the model. OpenAI says it fully disrupted that wider activity by July 28. A footnote on the 16,000 says the figures describe attempted extractions, not necessarily successful ones. An attempt is a request that tried the pattern. It is not, by itself, proof that a reasoning chain came back, and it is not proof that another model was trained on it. OpenAI says the activity changed over time, which is why it calls adversarial distillation a wider security problem that needs more than one layer of defense. Those figures are OpenAI’s.

Who OpenAI ties to the activity. The page says it is unclear whether every operator it observed in that period came from a single actor. It does attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. A core cluster is a central group inside the wider activity. It is not a claim that every one of the more than 15,000 users works for Moonshot. Associated with is OpenAI’s phrase. The post does not say Moonshot stole model weights, the files that store a model. It does not say a criminal charge was filed. Those lines are OpenAI’s.

Why OpenAI says the risk reaches past one campaign. Extracted reasoning could be used to train another model without keeping the safeguards that were applied to the original model’s answers. A safeguard, here, is a limit meant to stop a harmful or disallowed answer. At a large scale, OpenAI says, distillation can also move advanced capabilities to someone else without that someone else spending the same effort on safety. The concern grows, the post says, as models gain capabilities in dual-use domains. Dual use means a capability that can help and can also be misused. OpenAI says the risk is not unique to it, that similar techniques may affect other advanced systems, and that the defense has to be coordinated across the industry. Those lines are OpenAI’s assessment. They are not a measured copy of Kimi, and they are not a government finding.

What OpenAI says it did. It banned or restricted fraudulent accounts, strengthened the controls on signup and on its infrastructure, and expanded monitoring for related networks. It strengthened protections for hidden reasoning across users, workspaces, organizations, and model families. It closed a pathway that let someone who already had another user’s encrypted reasoning replay it and recover the contents. Replay means send that locked text back through a model. It added checks to detect and hold streamed output that might expose reasoning. Streamed output is the answer as it arrives, piece by piece, rather than all at once. When related activity moved through outside services, OpenAI says it worked with those providers to identify and disrupt the accounts. It shared findings through the Frontier Model Forum and through government information-sharing channels, so other developers and public-sector partners could look for similar activity and strengthen their own defenses. It says systems that support portable or replayable reasoning artifacts may face related risks. An artifact, here, is a saved piece of reasoning that can be carried from one place to another. Before publishing, OpenAI says, it looked at the scope and the possible impact, put its own mitigations in place, and shared the work with researchers and industry partners and took their feedback. It says more mitigation and investigation is still going on. Those lines are OpenAI’s.

What OpenAI says is still open. It expects these attempts to become more sophisticated as frontier models improve and as people look for cheaper ways to mimic what those models can do. Partner-hosted deployments, meaning copies running on a partner’s computers rather than only on OpenAI’s own service, need the same protections. Attacks that come through a tool’s output need checks that look at more than ordinary visible text. OpenAI says it is still improving those tool defenses, the classifiers that sort traffic, the cases where a model refuses a request, and the spread of those controls to cloud partners. A classifier is a check that sorts activity into categories. The response it describes from here has three parts: stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper sharing of threat information with industry and government. Those lines are OpenAI’s.

Bloomberg Law and The Verge reported the post the same Wednesday. Bloomberg Law’s headline is “OpenAI Blames Moonshot for Mass Data Extraction on Its AI Models.” The story, by Maggie Eastland, is stamped Sept. 30, 2026, 5:17 p.m. UTC. The opening says OpenAI accused its Chinese rival of a wide-scale effort to extract data from its GPT systems that could be used to reproduce the reasoning and capabilities of its most advanced models. The next sentences say OpenAI disclosed thousands of attempts by users associated with Moonshot to decipher hidden information about how its models reason, that the coordinated efforts began in early July, and that they peaked at 16,000 requests later in the month. Bloomberg also says OpenAI could not attribute all of the operators to a single actor. “Chinese rival,” “GPT,” and “users associated with Moonshot” are Bloomberg’s wording. OpenAI’s post says individuals associated with Moonshot AI for a core cluster, and it does not assign the whole 16,000-request spike to that cluster. The Verge, by Emma Roth, is posted September 30, 2026, at 6:56 p.m. UTC. Its headline is “OpenAI claims Moonshot extracted its data to train AI models.” The item says OpenAI, in a blog post on Wednesday, disrupted a coordinated campaign designed to extract protected reasoning, has not linked the activity to a single actor, and traced a core cluster to Moonshot, which The Verge calls the Chinese company behind the model Kimi. The Verge also writes that Anthropic has made similar accusations against Moonshot. That sentence is The Verge’s. It is not in OpenAI’s post. The Verge headline says data was extracted to train models. OpenAI’s footnote says the counts are attempts, not necessarily successful extractions.

The picture is OpenAI’s announcement graphic for this post. On a white field, the OpenAI mark sits at the top left. The large line reads “Disrupting a coordinated model-distillation campaign.” Under it, the graphic prints September 30, 2026. At the bottom left, the word OpenAI sits above openai.com. It is the company’s graphic.

In plain terms, OpenAI said on Wednesday that it found and stopped a coordinated effort to pull protected reasoning out of its models, activity it calls consistent with adversarial distillation, and that a core cluster of that activity is tied to people associated with Moonshot AI, which makes Kimi. The company says the operators did not break encryption, did not break into a database, and did not read stored chats. It says they steered conversations so hidden reasoning could be made visible, including by replaying encrypted reasoning from one chat into another. The counts it prints are attempts: 16,000 requests from over 4,000 users on July 24 and 25, a wider pattern across more than 15,000 users, and a disruption it dates to July 28. It says it banned accounts, closed the replay path, and shared the findings with other labs and with government. It does not say a model was cloned, a weight file was stolen, or a charge was filed.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources