AI

US Authorities Cite Systematic Distillation of US Models by Chinese AI Firms

Three US agencies cite industrial-scale distillation of advanced US models by Chinese AI firms, detailing DeepSeek and Moonshot methods.

8 min read Reviewed & edited by the SINGULISM Editorial Team

US Authorities Cite Systematic Distillation of US Models by Chinese AI Firms
Photo by Growtika on Unsplash

The National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the Federal Bureau of Investigation published a joint advisory by September 8, 2026. It stated that China-based artificial intelligence companies have carried out industrial-scale knowledge distillation against advanced U.S. models. The targets include ChatGPT, Claude, Gemini, and Grok. Since at least 2024, they allegedly spread out requests on the scale of millions to extract capabilities.

According to reporting by BeauHD for Slashdot, the activity was described as central to their development strategy, not as a partial supplement. The advisory characterized the Chinese companies’ methods as systematic extraction, an act of incorporating U.S.-owned features and capabilities at large scale.

China-based artificial intelligence companies are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core — not merely a supplement — of their AI development strategy

The above is a quotation from the advisory issued by the three agencies. The wording, core rather than supplement, is distinctive. It reflects the view that this was a strategic development method rather than isolated terms-of-service violations.

Large-Scale Distillation Flagged by Three

US Agencies

Knowledge distillation is originally a technique for transferring capabilities from a huge teacher model to a smaller student model. It uses the teacher’s outputs as synthetic data to teach reasoning, writing style, and patterns of problem-solving. In research, it is established as a means of improving efficiency. The issue lies in the terms of use for outputs and the scale. Many commercial models prohibit using their outputs to train competing models. The current allegation is that this prohibition was circumvented while operating at large scale.

The three agencies said multiple Chinese companies were involved. They allegedly distributed requests across numerous accounts and APIs, proxy connections, cloud providers, and third-party aggregators. The aim was apparently to make it difficult to identify the source of use and to evade defenses. The purposes also allegedly included bypassing geographic restrictions, terms of service, and safety measures built into advanced models.

According to CyberScoop reporting, specific company names and target models were cited. DeepSeek and Moonshot AI are said to be at the center. Both are suspected of using U.S. frontier models to generate synthetic training data. Calls routed through intermediaries were allegedly combined with distributed execution.

Moves by the Chinese Companies Named

DeepSeek allegedly used outputs from U.S. models to train R1 and R3. Four Claude variants, two Gemini variants, five ChatGPT variants, and Grok 4 were named. The targets are said to span a broad combination across generations. The resulting synthetic data was allegedly used to strengthen its own models’ capabilities.

The strengthened areas allegedly include agentic functions and optimization of question-answering. Creative writing and professional document creation were also targeted. This suggests an aim to improve versatility as products, not merely to imitate a single skill. The approach is seen as collecting outputs from multiple teachers to compensate for weaknesses.

For Moonshot AI, even broader extraction was alleged. It is suspected of distilling 18 U.S. models for use in training Kimi-K2 and Kimi K3. Fable 5, described as Anthropic’s currently most advanced commercial model, was also included. Agentic reasoning, code generation, and data analysis were among the extraction targets. Computer vision, large-scale reasoning frameworks, and visual processing were also cited.

The figure of 18 models indicates reliance beyond a single teacher. The method allegedly combines multiple strengths to raise performance for specific uses. Practical areas such as coding, analysis, and vision were selected. The intent to boost product competitiveness in a short period can be read from this.

Circumvention Methods Used to Evade Detection

The techniques cited in the advisory are said to form a multilayered structure for evading detection. First, there is distribution of requests. Queries are split across different accounts, models, and infrastructure to avoid highlighting abnormal volume. Second, there is concealment of routes. Official APIs, remote cloud providers, and third-party aggregators are used to dilute information about the source of use.

Third, there is the use of proxy connections and gray markets. Geographic restrictions are bypassed and obtained access routes are reused. This allegedly has the effect of delaying enforcement of terms of service and triggering of safety measures. Regarding the reality of API integration and the complexity of external routing, the aggregation-route issues discussed in Actual Integration Revealed by a Survey of 56 Business Systems on MCP and APIs partly overlap. There is no direct causation, but the structure in which more routes make auditing harder is common to both.

For cloud providers, the challenge is also serious. Leased computing resources may be used as relays for extraction. The boundary between legitimate and illicit use cannot be drawn by call volume alone. Text generation and code assistance also occur in large volumes for legitimate purposes. Both behavioral analysis and identity verification are required.

Even as on-device execution environments improve, leakage from the teacher side will not stop. Developments in open infrastructure were also covered in Mesa Rusticl Enables Mali Panfrost by Default and Surface Laptop 7th Edition Is the Best Buy, New Models Are Pricier. Apart from improving efficiency on the execution side, provenance management for training data remains an independent challenge. Measures on both fronts are needed.

What It Means That Distillation Is Central to

the Strategy

The advisory’s statement that this is the core, not a supplement, carries weight. It assesses that distillation was positioned as the main path of development rather than auxiliary efficiency improvement. It points to a structure that scales back proprietary pretraining and acquires capabilities from others’ outputs. While this can compress R&D costs and time, it normalizes intellectual property infringement.

The use of synthetic data itself is not unusual. Self-improvement, in which a company’s own models are trained on their own outputs, is also spreading. The problem lies in using other companies’ commercial outputs at large scale. This is likely to fall under competitive training prohibited by terms of service. There is a risk that even the behavior of models with safety alignment could be replicated.

Technically, it is difficult to identify traces of distillation. Writing style and reasoning habits may remain, but definitive attribution is said to be difficult. Mixing multiple teachers further dilutes distinctive features. Without records from the calling side and disclosure from the training side, verification cannot advance. The current allegation appears to be based on tracking of communications and procurement on the calling side.

Impact on US-China AI Competition and Future Focus

In the short term, monitoring by API providers is expected to intensify. Stricter anomaly-detection thresholds and expanded screening of aggregators will be the focus. Uses that send large volumes of queries in a short period will become subject to review. Monitoring will be stronger in high-value areas such as agentic functions and code generation.

In the medium term, countermeasures are expected to advance on both contractual and technical fronts. Revisions to terms of service, watermarking of outputs, and introduction of provenance management will be debated. Identity verification by cloud providers and linkage to payment information may also be strengthened. Excessive restrictions, however, would harm legitimate research and interoperability. Improved detection accuracy will be a prerequisite.

The geopolitical context cannot be ignored. Advanced U.S. models are positioned as central to national competitiveness. Capability leakage is treated as a national security concern. On this site over the past 30 days, the rise of Chinese models has also been a topic, including Kimi Files Confidential IPO in Hong Kong, Final Private Round Toward $50 Billion Valuation and Kimi API Now Natively Supports Codex and Claude Code. The current allegation will affect views of their rapid performance gains. This is a phase in which performance evaluation and provenance evaluation are required in parallel.

Editorial Opinion

We see this allegation tightening enforcement of terms of service by API providers. Detection of anomalous calls and countermeasures against anonymized use via aggregators could advance within three to six months. Identity verification and usage monitoring by U.S. cloud providers are also likely to be strengthened.

We assess that regulation of distillation will shift the boundary between open and closed models. Provenance management for synthetic data and disclosure requirements for training data could be institutionalized within one to three years. Developers will be pressed to secure proprietary data and ensure verifiability.

Where to draw the line between distillation and legitimate API use remains an open question. A blanket ban on using outputs for training could harm research and compatibility. The question is how to implement a verifiable provenance-labeling mechanism.

References

Frequently Asked Questions

What is knowledge distillation?
It is a technique for training a small model using the outputs of a huge teacher model as teaching material. It transfers reasoning steps, writing style, and patterns of problem-solving as synthetic data. While there are legitimate uses for efficiency, using outputs from another company's commercial models may constitute a terms-of-service violation.
Which US models and Chinese companies were targeted this time?
On the US side, ChatGPT, Claude, Gemini, and Grok were cited. On the Chinese side, DeepSeek and Moonshot AI were named. They are suspected of using the outputs to train DeepSeek's R1 and R3 and Moonshot's Kimi-K2 and Kimi K3.
What should API users watch out for?
They need to check contracts for conditions on reuse of outputs. Training competing models is prohibited under many commercial terms. Because the client remains responsible even when using aggregators, maintaining usage records and provenance management is required.
Source: Slashdot

Comments

← Back to Home