

A method that enables developers to compress advanced artificial intelligence models into more affordable and efficient systems has emerged as a new flashpoint in the escalating AI rivalry between the U.S. and China.
It is method that relies on the outputs of a large, powerful AI system to train a smaller model that can handle many of the same tasks while using fewer computing resources.
The most advanced AI systems, often called frontier models, demand vast computing resources, extensive datasets, and substantial financial investment to be developed.
Model distillation provides a method for building more compact systems by leveraging a large “teacher” model to train a smaller “student” model.
Once regarded as a core technique in AI research, distillation has now moved to the forefront of a growing controversy over whether advanced AI capabilities can be reproduced or passed on without the consent of the companies that first created them.
Washington and major U.S. AI companies have alleged that their Chinese competitors are employing this method to draw out functions from proprietary models, creating a new battleground in an increasingly acrimonious race for technological dominance.
The advantage of distillation lies in its ability to reduce the cost of AI and simplify its deployment.
A cutting-edge model might need massive data centers and costly processors to function. In contrast, distilled models can operate on more modest hardware and be customized for particular tasks..
This makes them attractive to corporations and public authorities seeking to roll out AI on a broader scale, spanning devices, industrial facilities, vehicles, and closed networks.
Distillation is a commonly used method for training AI systems and is not, by itself, an improper practice.
U.S. researchers and companies have been using it for a long time, among them Stanford University’s Alpaca initiative and Microsoft’s Orca project, both of which depended on the outputs of more sophisticated models to enhance the performance of smaller ones.
Chinese researchers have likewise incorporated results from U.S.-developed models into public research initiatives, including projects aimed at building Chinese-language instruction models.
The main distinction is access.
Open-weight models allow researchers to examine and adjust their internal parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, stay under corporate control and are usually available only through proprietary interfaces or APIs.
The main dispute concerns not the distillation process itself, but the unauthorised extraction.
AI companies contend that there is a clear difference between legitimate research and the systematic extraction of outputs from proprietary models to reproduce commercially valuable capabilities.
Anthropic has alleged that Chinese organizations, including DeepSeek, Moonshot, and MiniMax, have carried out extensive efforts to extract capabilities from Claude models.
The company said these initiatives focused on capabilities such as software engineering and advanced reasoning.
OpenAI has also reported identifying attempts by Chinese actors to use its models for distillation-related activities.
To date, no Chinese firms have publicly accused U.S. competitors of distilling closed-source models.