AI Policy10 min read

The Distillation Dilemma: When America's Spies Accused Six Chinese AI Labs of 'Stealing' Open Science

September 11, 2026·AI in China
The Distillation Dilemma: When America's Spies Accused Six Chinese AI Labs of 'Stealing' Open Science

heroImage: "https://images.unsplash.com/photo-1550751827-4bd374c3f58b?w=1200"

*Photo: Cybersecurity and digital sovereignty. When three US intelligence agencies jointly accuse six Chinese AI labs of systematic model extraction, the line between research technique and national security threat becomes the defining geopolitical question of the AI era. Image: Unsplash*


The Alert at Langley

It was 9:47 AM on Tuesday, September 8, 2026, when the joint advisory hit the wire. Designated AA26-251A, the document carried the logos of three of America's most powerful intelligence and security agencies—the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the Federal Bureau of Investigation. Its subject line was clinical, almost academic: *"China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against US AI Companies."*

Inside the language was something unprecedented. For the first time, US intelligence agencies had publicly named six specific Chinese artificial intelligence companies—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—and accused them of systematically extracting proprietary capabilities from American frontier models through a technique called knowledge distillation. The scale was staggering: millions of requests, billions of tokens extracted, and campaigns the agencies described as forming the "core—not merely a supplement" of these companies' AI development strategies.

The timing was not accidental. The advisory landed exactly one day before the Chinese Foreign Ministry's regular press conference, giving Beijing minimal preparation time. It came amid reports that the Trump administration was considering an outright ban on Chinese AI models. And it followed by less than five months the White House's NSTM-4 memorandum—itself a sweeping accusation that Chinese entities were conducting "deliberate, industrial-scale" distillation of US AI systems.

Within hours, the document had been downloaded by cybersecurity professionals, parsed by policy analysts, and debated by machine learning researchers around the world. The question it raised was deceptively simple: When does a standard machine learning technique become a national security threat?

Digital network connections

*Photo: Global digital networks. The distillation campaigns allegedly routed through API aggregators, cloud providers, and intermediary "transfer stations" spanning multiple jurisdictions. Image: Unsplash*


What the Advisory Actually Says

The technical claims in AA26-251A are specific, detailed, and—by the admission of the agencies themselves—partially circumstantial. The advisory alleges that since at least late 2024, the six named Chinese companies have engaged in systematic, large-scale distillation of US frontier models including variants of Anthropic's Claude, OpenAI's GPT, Google's Gemini, and SpaceX's Grok.

Distillation, as described by CISA in its announcement, is a legitimate machine learning technique in which a smaller or less capable model is trained using the outputs of a larger, more capable one. The technique is ubiquitous in AI research: Google's original 2015 paper on distillation has been cited over 16,000 times. OpenAI itself used distillation to create smaller, faster variants of GPT-4. Anthropic has published extensively on Constitutional AI, a distillation-adjacent technique.

What transformed distillation from research method to security threat, according to the advisory, was scale, intent, and evasion.

Allegation CategorySpecific ClaimTechnical Detail
Scale"Industrial-scale" campaignsMillions of requests extracting billions of tokens
Intent"Core" of development strategyDistillation not supplemental but central to model training
EvasionMulti-pathway routingNative APIs, remote cloud providers, third-party aggregators
Concealment"Transfer stations"API proxy gray markets to bypass regional restrictions
Account fraudFraudulent account poolsShared premium subscriptions, synthetic identities
Prompt injectionJailbreak techniquesExtracting hidden chain-of-thought (CoT) reasoning
Auto-switchingRapid path adaptationSwitching to new models within 24 hours of release

*Table 1: The seven categories of allegations in advisory AA26-251A. The agencies emphasize that the combination of scale, systematicity, and concealment distinguishes these campaigns from legitimate research distillation. Source: CISA AA26-251A.*

The advisory describes a sophisticated operational architecture. Chinese companies allegedly used "transfer stations"—API proxy services operating in legal gray zones—to route requests through jurisdictions without US export restrictions. They allegedly created pools of fraudulent accounts with shared payment credentials, distributed queries across IP ranges to avoid pattern detection, and employed prompt injection techniques to induce models to reveal their internal chain-of-thought reasoning—capabilities that US labs had deliberately hidden from users.

Most striking was the allegation that MiniMax pivoted to distilling a new Claude model within 24 hours of its release—a speed that suggests systematic monitoring and automated extraction infrastructure rather than ad hoc research.


The Named: Six Companies, Six Strategies

The advisory's naming of specific companies was itself a significant escalation. Previous US government statements on Chinese AI had been general; AA26-251A named names, attributed specific behaviors, and—in the case of some companies—described distinct extraction strategies.

CompanyDistillation TargetAlleged FocusStrategic Context
DeepSeekClaude, GPT, Gemini, GrokReasoning capabilities, specialized optimizations, domain-specific functionsBuilding general frontier models with limited domestic compute
Moonshot AI (Kimi)Claude Fable 5Training Kimi K3, particularly reasoning and agentic capabilitiesPreparing $50B Hong Kong IPO; K3 is 2.8T parameter open-source model
AlibabaClaude, GPT variantsQwen family improvement, cloud AI service capabilitiesCompeting with OpenAI/Anthropic in enterprise API market
MiniMaxClaude (rapid switching)General model capability extraction, 24-hour pivot to new releasesRecently IPO'd; building Talkie consumer AI platform
StepFunClaude Opus 4.1-4.5, GPT-5 familyStep 4 model programming and agent capabilitiesRising challenger with strong coding benchmarks
Z.AI (Zhipu)GPT-5.5, Claude Opus 4.8Chain-of-thought reasoning, GLM-5 series enhancementOpen-source leader; GLM-5.3-Flash topped OpenRouter charts

*Table 2: The six named Chinese AI companies and their alleged distillation targets. The diversity of targets reflects the breadth of China's AI ecosystem—no single model or company dominates. Source: CISA AA26-251A; company disclosures; industry reporting.*

The inclusion of Moonshot AI is particularly notable. The Beijing lab behind Kimi K3—a 2.8 trillion parameter open-source model that has shaken global AI markets—is reportedly preparing a $50 billion Hong Kong IPO. The advisory alleges that Moonshot distilled data from Claude Fable 5 to train Kimi K3, a claim the company has denied, attributing K3's performance to "underlying original architectural innovation."

For DeepSeek, the advisory challenges the company's public narrative. DeepSeek has marketed its models as achieving frontier performance with minimal compute—a claim that, if true, would represent a breakthrough in training efficiency. The advisory suggests a different explanation: that DeepSeek's performance gains were achieved not through architectural innovation alone, but through "distillation-generated synthetic data" that reduced the need for massive compute investments.


Beijing's Response: From Dismissal to Counter-Accusation

The Chinese government's response was swift, coordinated, and notably sharp in tone.

On September 9, Foreign Ministry spokesperson Mao Ning addressed the advisory at the ministry's regular press conference. Her statement was carefully calibrated: she did not engage with the technical specifics, but instead reframed the issue as a matter of principle.

"China's AI development is the result of high-level technological self-reliance and open cooperation. We have always believed that all parties should strengthen cooperation to promote open, inclusive, and beneficial AI development that serves the welfare of all humanity. We hope the US side will earnestly implement the important consensus reached by the two heads of state and refrain from making unfounded accusations and smears against China. Both China and the US are major AI powers and should strengthen cooperation."

The following day, China's Ministry of Commerce issued a stronger statement, calling the US accusations "groundless in fact and without legal basis" and characterizing the advisory as an effort to "politicize and weaponize" a normal technical and commercial practice.

Response LevelSpokesperson/BodyKey MessageTone
DiplomaticMao Ning, Foreign MinistryAI development is self-reliance + open cooperation; refrain from smearsMeasured, principled
TradeMinistry of CommerceAccusations are groundless; US is politicizing normal technique; double standardsSharp, accusatory
TechnicalIndustry commentatorsDistillation is standard ML; US companies do it too; advisory conflates technique with espionageAnalytical, defensive
AcademicChinese AI researchersOpen-source model outputs are public; training on public outputs is research normTechnical, legalistic

*Table 3: China's multi-layered response to advisory AA26-251A. The response strategy—diplomatic restraint from the Foreign Ministry, sharper pushback from Commerce, and technical rebuttals from industry and academia—reflects a coordinated but differentiated communications approach. Source: Chinese government statements, press reporting.*

The Commerce Ministry statement went further, accusing the US of "technological hegemony" and "compute monopoly"—framing the distillation accusation not as a cybersecurity issue but as a competitive tactic designed to prevent Chinese companies from accessing AI capabilities that US firms had already made publicly available through APIs.

This framing points to a genuine tension in the US position: the models being allegedly distilled were accessible through public APIs and consumer interfaces. The outputs being extracted were, by definition, generated in response to user queries. The boundary between "using a service" and "extracting its capabilities" is not technically sharp—and the US agencies acknowledged as much, noting that distillation itself is "a valid training method" that "can be misused."


The Technical Paradox: Standard Science or Systemic Theft?

The central tension in the distillation debate is that the technique being condemned is not merely legal—it is foundational to modern AI development.

Knowledge distillation was introduced by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean in a 2015 paper titled *"Distilling the Knowledge in a Neural Network."* The paper has become one of the most cited in machine learning. Its core insight—that a smaller "student" model can learn from the "soft targets" produced by a larger "teacher" model—underpins virtually every effort to deploy large AI models efficiently.

TechniqueDescriptionLegitimate UseAlleged Misuse
Knowledge DistillationTraining student model on teacher outputsCreating efficient model variants (GPT-4 → GPT-4o, Claude → Haiku)Systematic extraction of proprietary capabilities
Chain-of-Thought ExtractionPrompting model to reveal reasoning stepsResearch on interpretability, debuggingHarvesting hidden reasoning for competitor training
API AggregationPooling multiple model access pointsBuilding comparison tools, meta-modelsObfuscating distillation traffic
Prompt EngineeringCrafting inputs to elicit specific outputsApplication development, researchJailbreaking safety guardrails, extracting restricted capabilities
Synthetic Data GenerationUsing model outputs as training dataAugmenting limited datasets, privacy-preserving trainingBuilding competing models without original research

*Table 4: The techniques cited in AA26-251A and their dual-use nature. Every method described in the advisory has legitimate, widely practiced applications in AI research and development. The distinction between legitimate and illegitimate use depends on intent, scale, and contractual boundaries—none of which are technically precise. Source: Academic literature; CISA advisory; industry practice.*

OpenAI itself has built a significant portion of its product strategy around distillation. GPT-4o, the company's most widely deployed model, is understood to be a distilled variant of GPT-4. Anthropic's Haiku and Sonnet models are similarly positioned as efficiency-optimized derivatives of their larger Claude siblings. Google's Gemini family follows the same pattern.

What distinguishes these cases from the Chinese allegations is authorization and scale. OpenAI distills its own models. The Chinese companies, according to the advisory, distilled *others'* models—without authorization, in violation of terms of service, and at a scale that the agencies argue crosses from research into systematic extraction.

But this distinction is less clean than it appears. When a researcher at a Chinese lab uses Anthropic's public API to generate training data, is that research or extraction? When a startup buys API credits through a third-party reseller in Singapore, is that circumvention or legitimate market access? The advisory itself acknowledges ambiguity, using the phrase "likely with Chinese government awareness"—a carefully hedged formulation that stops short of claiming direct government orchestration.


The Geopolitical Context: An AI Cold War in Real Time

The distillation advisory did not emerge in a vacuum. It arrived at a moment of acute US-China tension over artificial intelligence—a tension that has been building since the Biden administration's chip export controls and accelerating under the Trump administration's second term.

TimelineEventSignificance
Oct 2022US restricts advanced chip exports to ChinaBegins hardware containment strategy
Oct 2023Biden expands chip controls to cover more countriesAttempts to prevent circumvention
Apr 2026White House NSTM-4 memo on "adversarial distillation"First government-wide framing of distillation as threat
May 2026Trump visits China; AI talks agreedDiplomatic opening on AI governance
Jul 2026Reports of planned Sept US-China AI talksNegotiated framework for AI risk discussion
Jul 2026Reports US considering ban on Chinese AI modelsKimi K3 open-source release triggers policy response
Aug 2026Gary Marcus: China AI "almost caught up" to USAcademic acknowledgment of capability convergence
Sep 8, 2026AA26-251A advisory issuedFirst naming of specific companies; most detailed public accusation
Sep 9, 2026China Foreign Ministry respondsDiplomatic pushback; reframes as cooperation issue

*Table 5: Timeline of US-China AI tensions leading to advisory AA26-251A. The advisory represents an escalation in a multi-year strategy of hardware containment, now expanding to include model-level restrictions. Source: Public reporting; government statements; academic commentary.*

The advisory also arrived alongside broader US concerns about Chinese AI capabilities. In July 2026, renowned AI researcher Gary Marcus wrote on his blog that Chinese AI models had "almost caught up" to American frontier models, and that the United States "cannot win" an AI zero-sum competition. Marcus urged Washington to "stop treating AI as a zero-sum game and instead explore international cooperation and public goods approaches."

The convergence of capability and open-source availability has created a policy dilemma that the distillation advisory attempts to address but does not resolve. If Chinese models like Kimi K3 (2.8 trillion parameters, fully open-source) and DeepSeek-V4 are genuinely competitive with American frontier models—and they are increasingly benchmarked as such—then restricting distillation of closed models may slow but cannot prevent Chinese AI advancement.

The advisory's recommended countermeasures reflect this reality. Rather than calling for sanctions or criminal charges, the agencies recommended detection, degradation, and information sharing: monitoring for anomalous usage patterns, subtly degrading responses to suspected distillation attempts, and sharing telemetry across model providers and cloud platforms.

Recommended CountermeasureDescriptionPractical Challenge
Anomaly DetectionMonitor subscription-to-usage ratios, peak usage, multi-IP accessNormal enterprise agent fleets exhibit identical patterns
Response DegradationServe suspected accounts weaker models without disclosureRisks degrading legitimate users; ethical concerns
Information SharingCross-vendor telemetry sharing on distillation campaignsCompetitive sensitivities; antitrust concerns
Terms EnforcementStrengthen API terms of service against distillationDifficult to enforce across jurisdictions
Geographic RestrictionsTighten regional access controlsVPNs and proxy services routinely circumvent

*Table 6: The five countermeasures recommended in AA26-251A. The agencies notably did not recommend sanctions, export controls, or criminal charges—suggesting recognition that the problem is not easily addressed through traditional enforcement mechanisms. Source: CISA AA26-251A.*


The Detection Problem: When Your CI Pipeline Looks Like Espionage

One of the most striking aspects of the advisory's technical recommendations is how closely its detection indicators describe normal enterprise behavior.

A widely circulated analysis by the Berkeley Existential Risk Initiative noted that "every detection indicator in this week's federal AI advisory describes a well-run enterprise agent fleet": sustained 24/7 usage with no idle periods; new subscriptions going straight to maximum throughput; one account hitting the API from many IPs; usage optimized for cache hits rather than task diversity.

"That is not a Chinese distillation ring. That is your CI pipeline, your ticket triage agent, and the FinOps work your platform team did last quarter."

This observation highlights a fundamental challenge: the boundary between legitimate high-volume API usage and malicious distillation is not technically detectable without access to intent and downstream use. A company running automated tests against an AI API may generate usage patterns identical to a distillation campaign. A research lab studying model behavior may issue prompts designed to probe reasoning capabilities—prompts that, to a detection system, look like attempted jailbreaks.

The advisory's recommendation to "subtly alter responses to suspected malicious distillation attempts" without informing users raises its own concerns. If implemented broadly, this would mean that AI systems might deliberately serve degraded outputs to users based on secret determinations of suspicious behavior—a practice with obvious risks for reliability, fairness, and accountability.


What This Means for the Global AI Ecosystem

The AA26-251A advisory is more than a cybersecurity alert. It is a statement of strategic intent: the United States now views the outputs of its frontier AI models as strategic assets warranting intelligence-agency protection, and it views systematic extraction of those outputs as a threat to national technological leadership.

This framing has implications that extend far beyond the six named Chinese companies.

StakeholderImpactLikely Response
US AI LabsIncreased pressure to detect and prevent distillation; potential liability for enabling extractionInvestment in usage monitoring; API restriction tightening; possible legal action against aggregators
Chinese AI CompaniesReputational risk; potential US market restrictions; increased compliance costsAccelerating domestic training infrastructure; diversifying data sources; open-source strategy emphasis
API AggregatorsRegulatory scrutiny; potential liability as "transfer stations"Geographic restrictions; enhanced KYC; possible exit from gray-market jurisdictions
Global DevelopersReduced API reliability; potential degradation of service; uncertain terms of useMigration to open-source models; self-hosting; regional API provider diversification
Open-Source CommunityIncreased relevance as alternative to restricted APIsAccelerated development of open-weight models; reduced dependence on commercial API providers
PolicymakersPressure to define boundaries between research and extractionPotential new regulations on model access; international AI governance frameworks

*Table 7: Ecosystem impacts of the distillation advisory. The advisory's effects will ripple through every layer of the AI stack, from model providers to end users. Source: Industry analysis; policy commentary.*

For Chinese AI companies, the advisory may paradoxically accelerate the trend toward self-reliance that US chip sanctions have already triggered. If distillation of American models becomes legally and practically risky, Chinese labs have strong incentives to invest more heavily in domestic training infrastructure, original data curation, and open-source model development—areas where they are already making significant progress.

The open-source dimension is particularly important. Models like Kimi K3 and DeepSeek-V4 are fully open-weight, meaning their parameters can be downloaded and run locally without API access. These models cannot be "distilled" in the same way as closed APIs because their weights are already public. If US restrictions on API access intensify, the competitive advantage may shift decisively toward open-weight models—many of which are Chinese.


The View from Both Sides of the Pacific

The distillation debate has generated intense discussion across Chinese and international social media, with perspectives ranging from technical defense to outright cynicism.

@硅谷观察员 (Silicon Valley Observer, Zhihu):

"蒸馏本身就是ML领域最基础的技术之一。OpenAI自己用GPT-4蒸馏出GPT-4o,Anthropic蒸馏出Claude Haiku,这叫优化。中国公司做同样的事,就叫'恶意蒸馏'。标准的定义取决于谁在做。"

*Translation: "Distillation is one of the most fundamental techniques in ML. OpenAI distilled GPT-4o from GPT-4, Anthropic distilled Claude Haiku—this is called optimization. When Chinese companies do the same thing, it's called 'malicious distillation.' The definition of standard depends on who's doing it."*

@AIResearchDaily (Twitter/X):

"The NSA advisory on Chinese model distillation is technically fascinating and politically predictable. What's striking is not that Chinese labs are distilling—everyone distills—but that they're apparently doing it at a scale and sophistication that triggered a three-agency joint response. That says more about Chinese AI capabilities than the advisory intended."

@科技评论老猫 (Tech Review Old Cat, Weibo):

"美方这份报告最有意思的地方在于,它实际上承认了六家中国AI公司的模型能力已经逼近甚至达到了美国前沿模型的水平。否则何必'蒸馏'?如果中国模型差很远,蒸馏再多也没用。这份报告是变相的认证。"

*Translation: "The most interesting thing about the US report is that it actually acknowledges the six Chinese AI companies' models have approached or reached the level of US frontier models. Otherwise why 'distill' at all? If Chinese models were far behind, no amount of distillation would help. This report is a backhanded certification."*

@ml_engineer_sarah (GitHub Discussion):

"As someone who works on model efficiency, the line between 'research distillation' and 'industrial-scale extraction' is genuinely blurry. If I run 10M prompts through Claude to study its reasoning patterns for an academic paper, that's research. If a company runs 10M prompts to train a competing model, that's extraction. The difference is intent and use, not technique. The advisory correctly identifies the problem but offers no technically viable solution."

@国际政治观察 (International Politics Observer, Douban):

"这不是技术问题,这是话语权问题。美国定义了什么是'合法'的AI研究,什么是'恶意'的技术使用。当你拥有最前沿的模型,你就可以把别人的追赶定义为盗窃。这跟当年美国指责中国'窃取'工业技术的套路一模一样。"

*Translation: "This is not a technical issue, it's a discourse issue. The US defines what constitutes 'legitimate' AI research and what constitutes 'malicious' technical use. When you own the most frontier models, you can define others' catch-up as theft. This is exactly the same playbook the US used when accusing China of 'stealing' industrial technology."*

@ frontier_ai_ethics (X/Twitter):

"The most under-discussed aspect of the distillation advisory: it recommends US AI companies serve *weaker models* to suspected users without telling them. This is a breathtakingly bad idea. Imagine a medical researcher using an AI for drug discovery, unknowingly receiving degraded outputs because their usage pattern 'looked suspicious.' The collateral damage would be enormous."

@AI产品经理小王 (AI Product Manager Little Wang, Xiaohongshu):

"我们公司做AI应用的,每个月API费用几十万。看了这份报告之后反而更想试试国产模型了。如果美国那边随时可能因为你的用量' suspicious'而降质服务,那商业风险太大了。至少DeepSeek和Kimi的API不会因为你用得多就偷偷给你换个弱模型。"

*Translation: "My company builds AI applications; we spend hundreds of thousands on APIs monthly. After reading this report, I'm more inclined to try domestic models. If the US side might degrade your service at any time because your usage looks 'suspicious,' the business risk is too high. At least DeepSeek and Kimi's APIs won't secretly swap you to a weaker model just because you use a lot."*


Conclusion: Drawing Lines in Sand

The AA26-251A advisory marks a watershed in the US-China AI competition—not because it reveals anything technically surprising, but because it makes explicit what had been implicit: the United States now views the outputs of its frontier AI models as strategic assets requiring intelligence-agency protection, and it considers systematic extraction of those outputs to be a threat comparable to intellectual property theft or cyber espionage.

The problem is that the technique being condemned—knowledge distillation—is not merely legal but foundational. It is how OpenAI builds GPT-4o. It is how Anthropic creates Claude Haiku. It is how the entire industry makes large models practical. The distinction between "optimization" and "extraction" depends not on the technique but on who performs it, at what scale, and with what authorization—boundaries that are political, not technical.

For China's AI industry, the advisory is both a challenge and an accelerant. The six named companies will face increased scrutiny, potential market restrictions, and reputational headwinds in Western markets. But the advisory also validates their progress: the US would not issue a three-agency joint alert against companies whose models were not genuinely competitive. And the recommended countermeasures—detection, degradation, information sharing—are defensive rather than punitive, suggesting that Washington recognizes the limitations of its ability to prevent distillation through enforcement.

The deeper question is whether the United States and China can find a framework for managing this competition that does not treat every research technique as a zero-sum battlefield. The two countries have agreed to hold AI talks in September 2026—a diplomatic opening that the advisory's timing seems designed to influence. Whether those talks can address the distillation dilemma, or whether the issue becomes another frozen conflict in a deepening technological cold war, will shape the future of AI development for years to come.

In the meantime, the API requests continue. Somewhere in Beijing, Hangzhou, and Shanghai, engineers are running prompts through Claude, GPT, and Gemini—some for legitimate research, some for product development, and some, according to the NSA, for systematic capability extraction. Distinguishing among them is a problem that no advisory has yet solved.


*Published: September 11, 2026 | Category: AI Policy | Reading time: ~16 minutes*

Sources: CISA Advisory AA26-251A; Defense One; The Paper (澎湃新闻); Caixin; China News Service; Ministry of Commerce of China; Ministry of Foreign Affairs of China; Ars Technica; The Neuron; Berkeley Existential Risk Initiative; Gigazine; EET-China; Reuters; Axios; Bloomberg; Gary Marcus personal blog; public company disclosures.

M

By Meeeeed

Editor at AI in China. Tracking Chinese AI companies, funding rounds, and the technologies reshaping global tech. More about me.

← Previous

Unitree's IPO Sprint: How a Chinese Robot Maker Went From Zero to ¥17 Billion in Eight Years

Next →

The Geneva of Silicon: Why the US-China AI Safety Dialogue Is Already a Victory for Beijing