According to a report in TechCrunch, artificial intelligence experts are casting doubt on claims that Chinese company Moonshot's large open-weight language model, Kimi K3, achieved its advanced capabilities solely through the distillation of the American company Anthropic's Fable model. The debate arises following allegations by U.S. officials regarding systematic copying and the use of banned chips.
U.S. Administration Allegations and Watermarks
Michael Kratsios, White House science advisor, stated that Moonshot, the company behind the development of the Kimi K3 model—currently the largest available open-weight LLM—built its model by copying Anthropic’s Fable model. According to Kratsios, the company did this while using chips that are not approved for export to China. "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," Kratsios wrote amid reported discussions regarding a ban on Chinese open-weight models, which are stirring controversy in the AI sector. Moonshot did not respond to questions regarding its training process, and Kratsios did not share further details regarding the sources of his allegations.
Kratsios's comments echoed remarks by Treasury Secretary Scott Bessent, who noted that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable." However, it is not clear what these watermarks consist of, and the U.S. Treasury Department did not respond to an inquiry on the matter.
Technological Doubts and Timelines Too Short
Despite the official allegations, AI experts are expressing significant skepticism that distillation—the process of querying a large language model to understand its inner workings and copy its capabilities—is the explanation for the advanced capabilities displayed by Kimi K3.
Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, explained to TechCrunch that timelines simply do not allow for it. "I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Hancock said. "There’s just not even frankly time, right? Fable's only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks."
Nathan Lambert, an AI researcher at the Allen Institute for AI, expressed a similar view in a recently released podcast. Lambert argued that the impact of distillation is decreasing as Chinese models get closer to the frontier of technology and the training regime shifts to reinforcement learning. According to him, "if it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning (SFT) alone."
Limitations of Fine-Tuning and the Need for Reinforcement Learning
Performing distillation requires a research lab to systematically query the target model to generate data that can be used in post-training stages. Sometimes this explicitly involves asking the model to explain its chain-of-thought to understand how it solves problems. In other cases, the prompts and responses of the model are used to train a new model in a process known as supervised fine-tuning (SFT).
This fine-tuning process is the reason why a model ostensibly built by a third party might claim during a conversation that it is Anthropic's Claude. According to Lambert, this is the stage where the model "picks up its manners." However, Lambert believes that the benefits of SFT are becoming less important as models become more complex.
To distill capabilities similar to those of Fable, reinforcement learning techniques would likely be required. In many cases, this means using an agent of the larger model to grade the smaller model's responses, and adjusting based on the grade given. The more advanced techniques also require highly significant infrastructure; large reinforcement learning runs can require tens of millions of agents. Using a frontier lab's API for this purpose would be insanely expensive and potentially a time bottleneck, because these models are pretty slow and, frankly, might not even provide a performance uplift.
History of Claims and Common Industry Practice
Nevertheless, it appears that previous frontier models might have contributed to Moonshot's developments. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models. Anthropic claimed it identified millions of exchanges between its models and users identified by IP addresses and metadata of these companies. These queries were described as distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use. Anthropic did not respond to queries from TechCrunch regarding Fable distillation.
At the same time, distillation is considered common among many AI companies, and not just in China. Elon Musk testified earlier this year that his company, SpaceXAI, distilled OpenAI models to develop Grok, adding that the practice was common in the industry. The line between distillation and developing synthetic datasets can be fairly blurry.
Hancock added that, in general, Americans are understating the technical expertise of these Chinese teams: "One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. If American models ground to a halt, I think China's progress would slow, but would still continue. They’re not just riding coattails here."
Chip Smuggling Routes and Data Center Oversight
It is hard to separate distillation claims from the second part of Kratsios's allegations—the claim that Moonshot obtained advanced Nvidia Grace Blackwell 300 chips (known as GB300), and also accessed servers equipped with GB300 chips in Thailand. The export of these chips to China is banned, but according to Sam Bresnick, a research fellow at Georgetown University's Center for Security and Emerging Technology (CSET), a black market exists for them.
In May, the founder of Supermicro, an American server builder, was indicted for smuggling advanced chips into China. Bresnick noted that he is a proponent of "know your customer" (KYC) laws for data centers across the world: "If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing."
President Joe Biden's Department of Commerce proposed federal "know your customer" rules for data centers in 2024, but no further progress appears to have been made under Donald Trump's administration. However, exporters shipping advanced chips abroad are required to ensure they are only used for approved purposes.