China's AI Models Are Beating American Ones on Some Tests. Is Distillation Really the Whole Story?
Anthropic and Washington say Chinese labs are stealing their way to the top. A growing number of experts say that story is incomplete, and possibly dangerous to believe.

Key points
- Cohere CEO Aidan Gomez, co-author of the foundational 2017 paper "Attention Is All You Need," says Chinese AI models now beat American ones on some benchmarks, which copying alone cannot explain.
- Anthropic's report this month names Alibaba and DeepSeek among companies that attempted "illicit distillation" of its Claude models.
- CISA described Chinese distillation campaigns as "industrial-scale" and called them the core of China's AI strategy, not a side element.
- US chip export controls may have forced Chinese developers toward leaner architectures, producing genuine innovation as a byproduct of constraint.
Distillation, in plain terms, is training a smaller AI model by feeding it answers produced by a more powerful one. The smaller model learns to mimic the bigger model's behaviour without the billions of dollars that built the original. In the US-China rivalry it has become a flashpoint: American labs call it theft; China's Ministry of Commerce calls the accusations "groundless and legally unsound."
Both things can be partly true.
What does the evidence actually show?
Chinese labs have demonstrably used distillation, and Gomez does not dispute that. What he disputes is that it explains everything.
"You can close the gap and reduce the gap by copying, but you can't outperform," Gomez told CNBC. His point is mechanical: a model trained entirely on another model's outputs cannot exceed the original. Beating it requires independent capability.
Some Chinese models now beat American equivalents on at least some benchmarks, standardised tests that measure what an AI system can and cannot do. That result, Gomez argues, is proof that genuine innovation is happening alongside whatever copying occurred.
Anthropic's position is harder-edged. Jacob Klein, the company's head of threat intelligence, told CNBC this month that "there's an entire illicit ecosystem to try to gain access to Claude and other models." AI2Day covered the broader federal case on 9 September, when the NSA, CISA and FBI named DeepSeek among six Chinese firms accused of extracting capabilities from US models since late 2024. CISA's framing then was the same as now: systematic, central, not incidental.
Sriram Krishnan, a former senior White House AI policy adviser, offered a different frame. Models like ChatGPT and Claude were themselves trained on human-generated content scraped from the internet, which is its own form of distillation. "The idea of distilling has always been a core part of how computer science works," Krishnan told CNBC's Squawk Box.
What does the chip story add?
China faces real hardware restrictions. US export controls have limited Beijing's access to Nvidia's advanced chips, the specialised processors that AI training depends on. That constraint has pushed Chinese developers to build leaner model architectures and squeeze more from cheaper hardware.
"Because of chip restrictions, Chinese developers were forced to build leaner, smarter architectures and do more with less," said Neil Shah, a partner at Counterpoint Research. "Dismissing that as mere imitation might make for convenient policy making, but it fundamentally misjudges the competition."
That matters beyond any single benchmark. A government backing AI this aggressively, while forced by sanctions to engineer around hardware limits, is not one coasting on stolen outputs. Gomez put it plainly to CNBC: "China is no longer just copying but is actually innovating outright."
Distillation is real and worth stopping. Treating it as the complete explanation for China's AI progress is a different claim entirely, and getting that wrong has policy consequences that outlast any test score.



