The Pulse
Anthropic Details Claude Misuse in Cyber, Bio and Distillation Campaigns
Anthropic says threat actors used Claude in cyberattacks, dual-use biological research and large-scale efforts to copy its model capabilities. The September 2026 report also identifies suspected Chinese distillation campaigns involving Alib

AI.info Team ·
Anthropic’s latest threat report begins with a problem that sounds less like chatbot misuse than a change in the operating model of cybercrime: human attackers set targets, while AI systems handle reconnaissance, tool development, exploitation, persistence and data theft. In one campaign attributed to a Russian state-linked actor, Claude-assisted workflows repeatedly rebuilt malware after security products detected it, then redeployed the modified tools against new victims.
The report, published September 10, covers activity Anthropic says it disrupted between December 2025 and August 2026. Its seven categories include cyber operations, surveillance, influence campaigns, fraud, biological misuse, conventional weapons development and illicit distillation—the covert extraction of one model’s capabilities to train another.
Anthropic says the cases involved its Haiku, Sonnet and Opus models. None of the misuse cases involved the company’s Fable or Mythos models, except for one distillation campaign. The report presents the incidents as selected examples rather than a complete census of misuse, and repeatedly separates evidence of suspicious activity from proof of malicious intent.
Claude Moves From Assistant to Cyber Operator
Anthropic’s most developed cyber case centers on GTG-20006, an actor the company says is consistent with public reporting about the Russian group Midnight Blizzard. The operation targeted military intelligence organizations in Ukraine and Europe, diplomatic and defense bodies, think tanks and people connected to U.S. foreign policy.
The group used Claude through customized workflows that automated much of the attack chain. Anthropic says the tooling covered infrastructure acquisition, phishing, command-and-control persistence, credential theft and data exfiltration. The actor also used AI to monitor whether malware had been detected; when a security product flagged a tool, automated agents modified and rebuilt it until it evaded the observed detection.
More than 20 organizations appeared in the actor’s operational planning and live activity, according to Anthropic. Targets included Ukrainian government and military personnel, drone manufacturers and suppliers, hotels that provided indirect access to travelers, a North African government technology authority, and diplomatic and government organizations using Microsoft 365.
The campaign did not depend on an unknown vulnerability or a novel form of malware. Anthropic’s finding is narrower and more consequential: familiar techniques such as phishing, stolen credentials, exposed services and DNS hijacking became cheaper to run across many targets because AI handled labor that once required multiple specialists.
“They’re effectively automating parts of the job within the intel apparatus,” Jacob Klein, Anthropic’s head of threat intelligence, told Axios. Klein said AI was making state-backed surveillance cheaper and more efficient without fundamentally changing whom governments target.
Five Biological Cases Expose the Dual-Use Problem
Anthropic’s biological findings are more cautious than the report’s cyber disclosures. The company does not claim that Claude enabled a completed biological weapon, and it does not identify the institutions, countries or specific agents involved. Instead, it describes five cases in which users sought assistance that could support dangerous biological research while also resembling legitimate scientific work.
One case involved a grant application for gain-of-function research on chikungunya, a mosquito-borne virus. Anthropic says the proposed work focused on transmissibility and immune evasion and sought to identify mutations, engineer them into infectious clones and select for increased virulence in animals. The company blocked the request and later found that the work was associated with a military research institute.
The users reached the model through a platform serving life-sciences researchers in regions where Anthropic does not officially provide service. According to the report, the platform routed traffic through U.S. infrastructure, used a zero-data-retention service and maintained fallback access to another model when Claude refused a request. Anthropic banned the linked accounts and worked with partners to disrupt the relay network, but says the operator rebuilt access within days using new identities and consumer subscriptions.
Another researcher used Claude for weeks while planning experiments related to how avian influenza viruses adapt to mammals. Anthropic says its classifiers confined the work to weaker models, limiting the assistance to clerical analysis, study design and related tasks. The company estimates that Claude’s contribution in that case was substantially below expert-level biological research.
Other cases were harder to classify. One user had Opus draft an orthopoxvirus grant application in about an hour, including the hypothesis, experimental design, dosing, statistical plans and contingencies. Two additional researchers developed computational systems involving venom peptides and toxins, framing the work around analgesics, antidepressants and other therapeutic applications while also creating outputs that could support harmful compounds.
Anthropic says one researcher deliberately kept the identities of a bacterial toxin and a hemorrhagic-fever virus protein vague in progress reports. The company banned the accounts for violating its supported-region policy, but did not assert that the researchers intended harm. A 30-day review of activity associated with state institutions found roughly 35 distinct research efforts, most of them ordinary civilian science, with some showing notable dual-use potential.
The company’s stated concern is not that every suspicious biological query represents a weapons program. Its concern is that the same technical information can support vaccines, treatments or outbreak research on one side and the design of more harmful biological agents on the other. Anthropic says filters can block explicit requests, but they cannot reliably infer intent from technically sophisticated work that appears beneficial when viewed one exchange at a time.
Alibaba, DeepSeek and Zhipu Target Claude’s Reasoning
Anthropic uses “illicit distillation” to describe industrial-scale campaigns that extract a model’s capabilities through large numbers of unauthorized requests and reproduce them in another system. Legitimate distillation is a standard training method: a smaller student model learns from responses generated by a larger teacher model. Anthropic says the difference in these cases lies in the covert scale, fraud and lack of authorization.
The company identifies seven China-based labs as responsible for additional distillation campaigns since February. It says the operations used fake identities, stolen payment cards, compromised API keys, proxy services and third-party model routers. Those services sometimes stored users’ conversations and sold them to other organizations, creating a separate privacy problem alongside the effort to copy Claude’s capabilities.
Anthropic calls an Alibaba operation its largest measured distillation campaign. The company says the campaign targeted reasoning transcripts from Opus 4.6 and 4.7, peaked at nearly 3 million exchanges per day and used more than 3,500 fraudulent accounts. The extracted material covered agentic tasks, software engineering, kernel development and long-horizon work, and was used to train or improve Alibaba’s Qwen models, according to Anthropic.
DeepSeek’s activity involved a different kind of exposure. Anthropic says DeepSeek routed selected requests from its own users through Claude, including conversations sent through coding tools such as Claude Code and other third-party harnesses. Over 14 days in July, the company observed more than 12.1 million exchanges attributed to DeepSeek. Some included internal corporate information, government data and live credentials, although Anthropic did not say that U.S. persons’ data was exposed in every case.
Zhipu, known outside China as Z.ai, used 273 fraudulent accounts in a ten-day operation against Claude Opus 4.8, Anthropic says. The company counted 770,609 exchanges passing through a chain-of-thought extraction system and attributed more than 3 million additional exchanges to Zhipu during the same period. Zhipu also used Claude to score outputs, clean transcripts, generate tasks and test models in its own training pipeline.
Anthropic separately attributes more than 400,000 exchanges across more than 1,500 accounts to Xiaomi. SenseTime allegedly purchased harvested Claude transcripts from data vendors, while Anthropic says MiniMax created a shell-company proxy service that offered access to Anthropic and OpenAI models but not Chinese models. The company says those findings suggest the service may have been built to collect exchanges from U.S. frontier models for model training.
Why Distillation Changes the Safety Equation
Anthropic says the danger is not limited to copying a coding or reasoning benchmark. General reasoning ability improves performance across many tasks, so a model trained on stolen outputs may gain capabilities beyond the subject matter of the harvested conversations.
In its own testing, Anthropic says a model distilled from a frontier system can acquire dangerous capabilities in biology and cyber operations even when the training exchanges contain little direct material about those fields. Safety controls applied to Claude do not automatically carry over to the unauthorized model that learns from Claude’s outputs.
The company also describes repeated attempts to extract internal reasoning through prompt manipulation. One lab tested more than 12,000 different requests to discover which techniques could reveal reasoning traces. Other attempts asked Claude to translate prior reasoning into another language or to treat a fake instruction as a system prompt.
Anthropic says it now uses metadata, irregular-activity signals and specialized classifiers to identify extraction campaigns. It has also changed how Claude presents internal reasoning, introduced safeguards that limit how new accounts can alter prompts and tools in multi-turn conversations, and moved toward broader account-level attribution rather than banning suspicious accounts one at a time.
Anthropic’s Preferred Answer Is Trusted Access
The report points toward a policy model that combines content filters with identity and institutional checks. Anthropic argues that high-risk biological capabilities cannot be offered safely through anonymous, general-purpose access because filters cannot distinguish all legitimate research from harmful work. Trusted programs, verified institutions and some degree of data retention would allow providers to investigate misuse while preserving access for approved scientists.
That approach creates a tradeoff the report does not resolve. More monitoring may help identify covert research and stolen model access, but it also means sensitive corporate, personal and government information can pass through intermediaries without users understanding where it goes. Anthropic’s examples include company financial forecasts, developer credentials and government-related material relayed from third-party model services.
The report’s most concrete message is that misuse is no longer confined to isolated jailbreaks. Cyber operators are assigning AI systems long sequences of tasks, researchers are using frontier models inside ambiguous biological programs, and competing labs are building industrial processes to copy model behavior. Anthropic says it banned the accounts it identified, shared intelligence with authorities and other AI companies, and strengthened its safeguards. The report also makes clear that shutting down one provider does not necessarily end an operation: actors can move to another model, a local deployment or a new proxy network.