Experts doubt China's Kimi K3 was built by copying Anthropic's AI model

White House officials claim Moonshot stole US technology to build its powerful new AI. Researchers say the timeline makes that story hard to believe.

AI2Day NewsdeskUpdated Editor: Lee Brown4 min read
Photoreal news-editorial 16:9 image of a modern smartphone lying face-up on a polished dark surface, its screen casting a faint glow upward, surrounded by trans
Share

Key points

  • White House science advisor Michael Kratsios accused Chinese AI company Moonshot of copying Anthropic's Fable model to build Kimi K3.
  • Anthropic released Fable on 1 July 2025, roughly two weeks before Kimi K3 appeared, a gap researchers say is too short for the copying technique alleged.
  • Treasury Secretary Scott Bessent said US government officials have found "watermarks" of American AI models inside Chinese ones, without specifying what those watermarks are.
  • Elon Musk testified earlier this year that his own company used the same copying technique on OpenAI models to build Grok, calling the practice common across the industry.
  • Separate allegations that Moonshot used Nvidia GB300 chips, which the US government bans from export to China, remain unaddressed by the company.

A senior White House official has accused a Chinese AI startup of stealing American technology. Researchers who study how AI models are built say the evidence is thin.

What exactly is the accusation?

Michael Kratsios, the White House science advisor, said Moonshot, the Chinese company behind Kimi K3, built the model through "distillation" of Anthropic's Fable model while using chips the US has banned from export to China. As we reported on 16 July, Kimi K3 was already drawing comparisons to Anthropic's most capable closed system before these accusations surfaced.

Distillation means systematically questioning an existing AI model to figure out how it thinks, then using those answers to train a cheaper copy. Think of it as reverse-engineering a recipe by tasting a dish repeatedly until you can reproduce it yourself.

Kratsios did not share the source of his allegations. Moonshot did not respond to questions about how it trained Kimi K3. Treasury Secretary Scott Bessent separately said officials have found "watermarks" of US models inside Chinese ones, but neither his office nor Anthropic explained what those watermarks actually are.

Why are researchers sceptical?

The timeline is the problem. Anthropic only made Fable publicly available on 1 July 2025. Kimi K3 appeared roughly two weeks later.

"You can't distill that much data and release it in two weeks," Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. He added that the industry tends to underestimate Chinese research teams: "One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work."

Nathan Lambert, an AI researcher at the Allen Institute for AI, made a separate point. Distillation has become less useful as Chinese models close the gap with US ones, because the technique works best when copying a much stronger model. At Kimi K3's level of capability, you would need reinforcement learning, where an AI grades its own answers repeatedly until it improves. Running that process through a commercial API would cost an enormous amount and might not even produce a performance gain.

Anthropoc did previously accuse Moonshot and DeepSeek of distilling its older models earlier this year, saying it found millions of unusual queries traceable to those companies' IP addresses, distinct from normal use. Whether any of that applies to Fable specifically is unanswered.

Distillation is not a China-only practice. Musk testified this year that his company used it on OpenAI's models to build Grok, calling it standard across the industry.

What about the chip allegation?

The second accusation, that Moonshot obtained Nvidia Grace Blackwell GB300 chips through banned export channels and used servers in Thailand, is a distinct concern. A black market for restricted chips does exist. In May 2025, the founder of US server manufacturer Supermicro was indicted for smuggling advanced chips into China.

Sam Bresnick at Georgetown's Center for Security and Emerging Technology argues that data centres worldwide need clear rules about who is running large training jobs on their hardware. A Biden-era proposal for such "know your customer" rules has not advanced under the current administration.

The broader political context matters here. Our earlier story on the fight over Chinese open-source AI found that the push to restrict models like Kimi K3 is as much about protecting commercial interests as about security. The theft accusation fits that pattern: it lands just as Washington is debating whether to ban Chinese open-weight models outright.

Common questions

Does distillation always count as theft?

Not straightforwardly. The line between distillation and building a synthetic dataset is blurry. US companies, including those with prominent founders, use similar techniques routinely.

Should I worry that Chinese AI models are less trustworthy as a result?

This dispute is about how models are trained, not how they behave for users. Questions about data privacy and security in AI tools are worth asking of any model, regardless of its country of origin.

© 2026 AI2Day