What is the story about?
Operators affiliated with Alibaba ran the largest illicit AI distillation campaign Anthropic has ever detected, extracting the reasoning of its Claude models at a peak rate of nearly 3 million exchanges a day to help train the company's Qwen models, Anthropic said in its latest threat report.
According to Anthropic's Detecting and Countering Misuse of AI: September 2026 report, the operation targeted the chain-of-thought reasoning transcripts of Claude Opus 4.6 and 4.7.
It recorded more than 151 million exchanges attributable to Alibaba between May and July 2026. Anthropic calls it the largest such attack it has "ever measured."
How it worked
The pipeline injected a fixed prompt into every request, forcing Claude to write out its reasoning inside inline text tags before giving its final answer, according to the report. Those transcripts were then converted into supervised fine-tuning (SFT) data used to help train Qwen 3.5, 3.6 and 3.7.
The attacks concentrated on agentic tasks, software engineering, kernel development and long-horizon reasoning, and peaked at nearly 3 million exchanges a day from more than 3,500 fraudulent accounts.
Anthropic says Alibaba's use of Claude extended beyond harvesting reasoning traces: the company also used Claude to help build its internal model-development infrastructure, its reinforcement-learning environments, and to advance model-architecture research.
Two waves of fake accounts
The report says Alibaba accessed Claude through two pools of fraudulent accounts. The first, roughly 5,000 accounts, used residential proxies, disposable emails and virtual-card payments to obscure its origin.
When Anthropic banned that pool, traffic shifted to a second one — some of which, Anthropic says, was also found funnelling requests on behalf of DeepSeek and Xiaomi, indicating the same proxy networks are shared across multiple Chinese AI firms.
Anthropic defines illicit distillation as industrial-scale, unauthorised extraction of a rival model's capabilities, usually enabled by fraud — fake accounts opened with stolen credit cards, credentials or API keys.
It says the risk extends beyond intellectual property: a distilled model can pick up capability gains "across tasks and domains, not just those targeted by distillation attacks," including in sensitive areas, without inheriting the original model's safety guardrails.
More labs running such campaigns
Alibaba is one of seven China-based labs Anthropic says it has identified running such campaigns against Claude since first disclosing the practice in February 2026; the report also names DeepSeek, Moonshot AI, Xiaomi, Zhipu (Z.ai), SenseTime and MiniMax.
None of the campaigns succeeded against Anthropic's Mythos-class models, which are not publicly accessible, the company claims.
Anthropic, in response, said that it has built classifiers to detect adversarial extraction, moved to banning accounts by attributed organisation rather than individually, and now requires identity verification for accounts showing signs of abuse or originating from unsupported countries including China, Russia and Iran.
It has also made Claude summarise its reasoning before responding, and introduced "preserved thinking" in its Fable 5.1 model, which stops new API accounts from altering the conversation context that precedes Claude's reasoning — a common tactic used to expose it.
The allegations against Alibaba come solely from Anthropic's own investigation; they have not been independently corroborated, and Anthropic does not disclose the technical basis for attributing the campaign to Alibaba. Alibaba had not publicly responded to the report at the time of writing.
According to Anthropic's Detecting and Countering Misuse of AI: September 2026 report, the operation targeted the chain-of-thought reasoning transcripts of Claude Opus 4.6 and 4.7.
It recorded more than 151 million exchanges attributable to Alibaba between May and July 2026. Anthropic calls it the largest such attack it has "ever measured."
How it worked
The pipeline injected a fixed prompt into every request, forcing Claude to write out its reasoning inside inline text tags before giving its final answer, according to the report. Those transcripts were then converted into supervised fine-tuning (SFT) data used to help train Qwen 3.5, 3.6 and 3.7.
The attacks concentrated on agentic tasks, software engineering, kernel development and long-horizon reasoning, and peaked at nearly 3 million exchanges a day from more than 3,500 fraudulent accounts.
Anthropic says Alibaba's use of Claude extended beyond harvesting reasoning traces: the company also used Claude to help build its internal model-development infrastructure, its reinforcement-learning environments, and to advance model-architecture research.
Two waves of fake accounts
The report says Alibaba accessed Claude through two pools of fraudulent accounts. The first, roughly 5,000 accounts, used residential proxies, disposable emails and virtual-card payments to obscure its origin.
When Anthropic banned that pool, traffic shifted to a second one — some of which, Anthropic says, was also found funnelling requests on behalf of DeepSeek and Xiaomi, indicating the same proxy networks are shared across multiple Chinese AI firms.
Anthropic defines illicit distillation as industrial-scale, unauthorised extraction of a rival model's capabilities, usually enabled by fraud — fake accounts opened with stolen credit cards, credentials or API keys.
It says the risk extends beyond intellectual property: a distilled model can pick up capability gains "across tasks and domains, not just those targeted by distillation attacks," including in sensitive areas, without inheriting the original model's safety guardrails.
More labs running such campaigns
Alibaba is one of seven China-based labs Anthropic says it has identified running such campaigns against Claude since first disclosing the practice in February 2026; the report also names DeepSeek, Moonshot AI, Xiaomi, Zhipu (Z.ai), SenseTime and MiniMax.
None of the campaigns succeeded against Anthropic's Mythos-class models, which are not publicly accessible, the company claims.
Anthropic, in response, said that it has built classifiers to detect adversarial extraction, moved to banning accounts by attributed organisation rather than individually, and now requires identity verification for accounts showing signs of abuse or originating from unsupported countries including China, Russia and Iran.
It has also made Claude summarise its reasoning before responding, and introduced "preserved thinking" in its Fable 5.1 model, which stops new API accounts from altering the conversation context that precedes Claude's reasoning — a common tactic used to expose it.
The allegations against Alibaba come solely from Anthropic's own investigation; they have not been independently corroborated, and Anthropic does not disclose the technical basis for attributing the campaign to Alibaba. Alibaba had not publicly responded to the report at the time of writing.
/images/ppid_59c68470-image-178910758822286021.webp)

/images/ppid_59c68470-image-178910752498742699.webp)
/images/ppid_59c68470-image-178911003610995635.webp)







/images/ppid_59c68470-image-178909512057616452.webp)
/images/ppid_59c68470-image-178909266069320087.webp)
