Washington accuses six Chinese AI firms of mass model theft

iEXExchanger
Washington accuses six Chinese AI firms of mass model theft

NSA, CISA and the FBI issued a rare joint advisory naming DeepSeek, Alibaba, Moonshot AI and three other firms for industrial-scale distillation of Claude, GPT, Gemini and Grok since late 2024.

Three US security agencies just did something they rarely do outside hacking gangs and ransomware crews: they put out a joint warning naming private companies. On September 8, the NSA, CISA and the FBI issued a shared advisory calling out six Chinese AI firms — DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI.

The accusation centers on distillation, a technique where a smaller model is trained to mimic a larger one by feeding it thousands of that model's own answers. It's common and mostly legitimate — the issue here is scale and secrecy. The agencies say the six firms pulled billions of tokens across millions of queries from American models since late 2024, treating distillation not as a side experiment but as the core of their development strategy.

The targeted systems span Claude from version 3.7 through Opus 4.8, including Fable, GPT from version 4 through 5.5 plus GPT-4o specifically, Gemini 2.5 and Grok. The advisory gets specific: DeepSeek allegedly drew on Claude, Gemini, GPT and Grok outputs to train R1 and V3, while Moonshot AI reportedly used Claude Fable data for Kimi K3 and GPT-4o data for Kimi K2.

  • Fraudulent accounts with scrubbed metadata
  • Proxy "transfer stations" to dodge geographic blocks
  • Bulk premium subscriptions shared across dev teams
  • Prompt injection and jailbreaking to bypass filters
  • Centralized routing that masked the true source of requests

None of this appeared out of nowhere. Back in June, reports surfaced that Alibaba trained Qwen on 28.8 million Claude responses in just 45 days. In July, the Treasury was reportedly weighing sanctions against Moonshot AI over similar claims. Until now, those were separate, company-by-company stories. A joint advisory from three agencies at once reads as something different — an official position that this is a pattern across the industry, not an isolated case.

Notably, the recommendation isn't "block them." The agencies suggest something subtler: quietly degrading response quality for accounts flagged with high confidence, rather than banning them outright, since outright bans just push offenders to spin up new accounts and keep going at the same pace. Worth remembering this is an advisory, not a verdict — no sanctions have been announced, and none of the named companies has issued a public response so far.

The real question is whether this warning turns into something with teeth — the kind of sanctions floated against Moonshot AI back in July — or stays a advisory with no practical follow-through.

Questions and answers

Frequently asked questions about this article

What is AI model distillation, and why is it a problem here?

It's a training method where a smaller model learns to mimic a larger one using that model's own answers as training data. The technique itself is legal, but here the agencies say it was done covertly, at industrial scale, and by evading service restrictions through fake accounts and proxy networks.

Which US AI models were affected?

According to the advisory, the targeted systems were Anthropic's Claude (versions 3.7 through Opus 4.8, including Fable), OpenAI's GPT (versions 4 through 5.5, plus GPT-4o specifically), Google's Gemini 2.5 and xAI's Grok.

Have sanctions already been announced against the named companies?

No. The NSA, CISA and FBI document is an advisory with recommendations for US companies, not a legal or regulatory ruling. No sanctions have been announced yet, though the Treasury reportedly considered exactly that step against Moonshot AI back in July over a similar issue.

How are companies like Anthropic and OpenAI advised to respond?

The agencies recommend not banning suspicious accounts outright but quietly degrading response quality for those flagged with high confidence, while also monitoring subscription-to-usage ratios and sharing threat intelligence across the industry.