AI Model Distillation’s Controversy and Facts

Distillation

Controversy surrounding distillation

Competitors Steal by Distillation

A major concern surrounding open-source AI is model distillation, where smaller models are trained using the output of larger models. Recent high-profile cases include allegations that Chinese developers Alibaba and Moonshot AI distilled Anthropic’s Claude model and used it to build their own systems.

Distillation is Inevitable used by all vendors

In the AI ​​industry, distillation is a basic technique that is difficult to avoid. Most importantly, all AI companies worldwide have been using it; none are inherently superior.

All manufacturers inevitably use distillation technology; no one is an exception! Those who oppose distillation, advocate for others to use it illegally, or even advocate for banning Chinese counterparts to establish their own moral high ground cannot withstand scrutiny—this is why the tide has turned, and American tech companies are now rising up against the government; because they understand this better than anyone else.

Western companies have long been using distillation

Distillation isn’t unique to Chinese companies; this technique has long been widely adopted by many Western tech companies. Now, the US opposes American companies using Chinese AI models, accusing Chinese AI companies of using distillation to steal their research.

Anthropic, OpenAI, and Alphabet were the three most vocal opponents of distillation. However, the tide has turned: Anthropic is now the only publicly and consistently accusing Chinese companies of distilling their work.

OpenAI, which previously criticized China, has now fallen silent. Alphabet discovered that its subsidiary DeepMind was actually the originator of distillation and has therefore chosen to remain silent.

Mistral, SpaceX AI, Meta, and Microsoft have been using distillation to enhance and improve their own AI models.

In an interview with CNBC on August 3, Palantir CEO Karp stated, “Someone is trying to get us addicted, as if we can only rely on them (referring to cutting-edge AI models like Anthropic) to control the future.” Karp said he doesn’t see anything unfair about China using model distillation to replicate American AI models, because cutting-edge AI models are essentially doing the same thing. Karp said, “Where do you think the value of these models comes from? They distill the value of IP from all over the world, including corporate data, almost everything. We are in competition, and these technologies have to really make a difference.”

In May 2026, Musk admitted in court that SpaceX.AI used OpenAI to train Grok for AI distillation, sparking controversy over double standards within the industry. He did not deny it, but responded that “generally speaking, all AI companies do this,” and further admitted that “some of them do.”

Western Also Distill Chinese AI Companies

In fact, many Western AI companies distill Chinese models during research and training. In 2025, Mistral was exposed for distilling DeepSeek models, plunging into a public relations crisis: the technical community discovered that some of its models were highly similar to DeepSeek in their generation style, and a former Mistral employee revealed that the company deliberately concealed the distillation process, misleading the results into being presented as self-developed technology.

In April 2026, Meta also stated that Muse Spark’s training used several third-party open-source models, including Qwen from Chinese tech giant Alibaba, as well as models from OpenAI and Google. However, the practice of using Chinese models clearly goes against the stance of some US policymakers and high-level technologists. In response, a Meta spokesperson stated, “Like other companies in the industry, Meta uses techniques such as filtering and refining to learn from publicly available AI models under strict safeguards to improve our own models.”

Thinking Machines’ first model, Inkling, primarily uses DeepSeek-V3 in its hybrid expert architecture. Its cold start training also utilized data from open models such as K2.5 from the Chinese AI startup Moonshot AI.

Cursor’s Composer 2 coding model is built directly on Moonshot AI’s Kimi K2.5 through a licensing partnership, and overlaid with Cursor’s own training data. Cursor acknowledges that its Composer 2 coding model is built on Moonshot AI’s Kimi K2.5 through a licensing partnership, and then overlaid with its own training data.

Anthropic distilled as well

Anthropic itself built its model through a similar distillation process. Mark Suman, CEO of AI startup Maple, stated, “Anthropic distilled the entire network without paying royalties, and now they’re asking others not to do what they did—it’s quite strange.”

Chinese ByteDance denies AI distillation

The US has accused mainland Chinese models of learning from US AI models through “distillation.” On August 6, 2026, Zhang Yiming, founder of ByteDance, the parent company of TikTok, publicly refuted these unfounded accusations. He unusually commented on the company’s model development roadmap, emphasizing that the company “rejects distillation” and believes that model development should adhere to long-term principles, sacrificing some short-term gains for long-term goals.

It is understood that ByteDance strictly prohibits the distillation of open-source models internally, even strengthening related restrictions internally through API testing and other methods. Zhang Yiming believes that distillation would interfere with true long-term technological breakthroughs.

The Information, a US technology media outlet, mentioned that ByteDance employees stated that ByteDance’s unwillingness to use distillation is considered one of the reasons why its large-scale language models lag behind other AI labs in mainland China. However, the report also points out that ByteDance is the only major AI company in mainland China that uses proprietary models (closed-source models) for almost all large-scale language models, while other companies in the industry mostly focus on open-source models.

Distillation

Related articles

Disclaimer

  • The content of this site is the author’s personal opinions and is for reference only. I am not responsible for the correctness, opinions, and immediacy of the content and information of the article. Readers must make their own judgments.
  • I shall not be liable for any damages or other legal liabilities for the direct or indirect losses caused by the readers’ direct or indirect reliance on and reference to the information on this site, or all the responsibilities arising therefrom, as a result of any investment behavior.
error: Content is protected !!