The Pulse
US Agencies Accuse Six Chinese AI Firms of Model Distillation
The NSA, FBI and CISA say six China-based AI companies extracted billions of tokens from U.S. frontier models through distributed, industrial-scale campaigns. The agencies urge model providers, cloud companies and API aggregators to share s

AI.info Team ·
Washington has turned a long-running dispute over AI copying into a formal cybersecurity warning. The National Security Agency, FBI and Cybersecurity and Infrastructure Security Agency say six China-based companies have extracted billions of tokens from American frontier models through millions of exchanges, then used those outputs to train competing systems.
Beijing rejects the accusation. Mao Ning, a spokesperson for China’s Ministry of Foreign Affairs, told a regular news conference on September 9 that “China’s AI development is the result of high-level technological self-reliance and strength,” and that “all parties should strengthen cooperation to promote AI development that is open, inclusive, universally beneficial and oriented toward the common good.” The ministry urged Washington to refrain from unfounded accusations; Mao said she had not seen the advisory itself. The clash leaves AI companies facing a practical question: how can they block organized extraction without shutting down legitimate research, commercial use and international access?
The agencies’ September 8 joint cybersecurity advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. It says the companies targeted versions of Anthropic’s Claude, OpenAI’s GPT models, Google’s Gemini and xAI’s Grok, with activity dating back at least to late 2024.
The agencies describe distillation as an organized supply chain
Knowledge distillation is a standard method in machine learning. A smaller “student” model learns from the outputs of a larger “teacher” model, allowing developers to reproduce selected capabilities with less computing power and less training expense. The U.S. advisory does not characterize the technique itself as improper; it distinguishes ordinary distillation from what it calls aggressive, malicious and targeted extraction of restricted capabilities.
According to the NSA, FBI and CISA, the named companies distributed their activity across model providers, cloud platforms, API aggregators, proxy services and account pools. That structure was designed to prevent any single provider from seeing the full campaign. The advisory says some operators used gray-market API proxies, called “transfer stations,” to bypass geographic restrictions, conceal the origin of requests and weaken traceability.
The report says operators also bought premium subscriptions in bulk, shared access across teams and rotated accounts when providers blocked traffic. Detection indicators include accounts that begin using their full quota immediately, shared credentials appearing from multiple IP addresses, uninterrupted activity around the clock and similar prompts arriving across different accounts.
Such behavior matters because model providers often review abuse one account or API key at a time. A distributed campaign can look like a collection of unrelated users unless providers compare timing, prompts, payment information, routing infrastructure and usage patterns. The agencies want those signals combined across the companies that sit between a user and a model.
DeepSeek and Moonshot drew the report’s most detailed allegations
The advisory gives its longest descriptions to DeepSeek and Moonshot AI. It says DeepSeek ran an organized campaign from at least late 2024 to generate synthetic training data for models including R1, released in early 2025, and V3. The agency report says DeepSeek targeted reasoning, legal specialization, question-and-answer performance, agent functions, writing and other domain-specific behavior.
The document lists a wide range of models that DeepSeek allegedly queried, including Claude Sonnet and Opus variants, Google’s Gemini systems, GPT-4, GPT-4o, GPT-5 and Grok 4. It also says DeepSeek used prompts designed to elicit step-by-step reasoning from models that normally restrict visibility into internal reasoning traces. The resulting material could teach a student model methods for coding, logical proofs and complex agent tasks rather than merely transferring factual answers.
Moonshot AI, the developer of the Kimi family, allegedly ran a separate campaign beginning at least in mid-2025. The advisory says the company extracted Claude Fable 5 data for Kimi-K3 and GPT-4o data for Kimi-K2, while also targeting software engineering, mathematics, reinforcement learning and supervised fine-tuning performance.
U.S. officials also name Alibaba, MiniMax, StepFun and Z.AI. The report says Alibaba used Claude and GPT-5 outputs to improve software engineering, customer-service dialogue, virtual-character creation and agent workflows. MiniMax allegedly targeted Claude Code and other models for coding, reasoning and dataset refinement, while StepFun focused on coding and agent functions. The advisory attributes billions of GPT-5.5 and Claude Opus 4.8 tokens to Z.AI for work on chain-of-thought reasoning.
“Transfer stations” let operators spread the risk
The technical details show why the agencies treat the activity as more than ordinary scraping. Operators allegedly routed requests through native APIs, remote cloud providers, third-party aggregators, relays and pools of vendor accounts. Central routing systems could select the cheapest available pathway, monitor whether a provider was responding, enforce quotas and switch providers when one route failed.
The advisory describes four techniques that the agencies say are not fully covered by the existing MITRE ATLAS framework for AI threats. They include regional-restriction evasion and subscription exploitation, centralized request routing, automated removal of identifying metadata, and systematic quota and cost optimization.
Prompt engineering played a role as well. The report says operators used jailbreaks and prompt injections to force models to reveal hidden reasoning or produce responses in formats useful for training. One example says DeepSeek prompted models to imagine and write out the internal reasoning behind completed answers. Another says MiniMax redirected exchanges to a newly released Claude model within 24 hours, suggesting that infrastructure was ready to switch targets quickly.
Campaigns allegedly ran for days or months, with thousands to millions of queries directed at particular domains. Moonshot is described as moving beyond text responses into agentic reasoning, tool use, computer-use agents and computer vision. The agencies say the goal was not indiscriminate collection, but systematic harvesting of capabilities that could be incorporated into competing models.
The advisory presents attribution as a U.S. government assessment, not a court finding. The named companies did not receive an adjudicated finding of liability in the document.
Washington warns that chip controls do not close the access gap
The report frames model distillation as a national-security issue because it can reduce the cost and time required to develop advanced systems. Chinese developers can obtain selected behaviors through APIs without building the same training infrastructure, buying the same volume of accelerators or repeating the same foundational research. U.S. officials say those capabilities could strengthen military, cyber and critical-infrastructure applications.
That argument shifts attention from physical computing hardware to the software interfaces through which models deliver their capabilities. Export controls can restrict access to advanced chips, but they do not automatically prevent a determined operator from querying a model hosted elsewhere. An API, a cloud reseller or a proxy service can become a channel for transferring behavior even when the underlying model never leaves the provider’s servers.
The advisory says the activity also creates direct financial harm for American AI companies. Providers pay for computing, research staff, data, safety testing and infrastructure before charging customers for access. A rival that copies selected capabilities through unauthorized requests may reduce its own development costs while using the original provider’s systems as an unpaid research pipeline.
China’s Foreign Ministry disputes that account. In remarks published by the ministry on September 9, Mao said China’s AI development reflects “open cooperation” and urged both countries to cooperate rather than trade accusations. No response from the six named companies appears in the NSA advisory.
Providers are told to change model behavior, not just block accounts
The agencies recommend three immediate measures. Providers should detect anomalous prompts, accounts, networks and behaviors; make targeted changes to responses from high-confidence malicious requests; and share intelligence across model companies, cloud platforms and API aggregators.
Changing responses is a more aggressive step than suspending accounts. The advisory says providers could subtly degrade the value of outputs sent to suspected distillation operations, reducing the quality of material available for training without necessarily revealing the detection system. Such measures would require careful controls because a mistaken classification could harm an ordinary customer, researcher or developer.
Cross-company coordination presents another difficulty. Model providers may be reluctant to share telemetry because of privacy obligations, competitive concerns or legal restrictions. The agencies nevertheless argue that distributed activity cannot be identified reliably by examining one provider’s records in isolation.
The advisory recommends watching for linked accounts, synchronized activity, repeated prompt templates, rapid switching among models and usage patterns that prioritize volume over task diversity. It also points to metadata changes after public disclosures as a possible sign that an operator has altered its infrastructure to avoid detection.
The immediate result is a new defensive burden for companies that sell access to frontier models. They must decide which users deserve heightened scrutiny, how much behavior they can alter, and when evidence is strong enough to share with another provider or a government agency. The September 8 advisory supplies a framework for those decisions, but it does not settle the question of how much access providers should continue to offer while the U.S.-China dispute moves through commercial APIs.