Anthropic on September 10, 2026 published its most detailed threat intelligence report to date, and the findings read like a rundown of every way a frontier AI model can be turned against its own safeguards. State-linked operators in Iran, Russia and China, along with Yemen-based groups tied to Houthi militants, tried to bend Claude toward weapons research, propaganda, surveillance and cyberattacks over an eight-month stretch. Anthropic says it caught and shut down the activity, but the report itself marks a shift: models the company once considered too weak to meaningfully assist with weapons of mass destruction are now, in its own assessment, past that line.
The report, titled “Detecting and Countering Misuse of AI: September 2026,” covers activity Anthropic disrupted between December 2025 and August 2026. It lands alongside a separate NewsGuard and NPR test of six AI chatbots that found the bots still repeat state propaganda a meaningful share of the time. Together, the two findings give the clearest public picture yet of how Tehran, Moscow, Beijing and non-state armed groups are actually using commercial AI systems in 2026, not how analysts speculate they might.
What Anthropic’s September 2026 Report Actually Says
Anthropic’s report groups misuse into seven categories: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development and illicit model distillation. The activity spans Claude’s Haiku, Sonnet and Opus model lines. Anthropic notes that none of the confirmed misuse cases involved its newer Claude Fable or Mythos-class models, with one exception in the distillation category, according to the company’s own report.
Reuters, which reviewed the report ahead of publication, wrote that Anthropic said several actors had used its Claude models “for activities ranging from weapons development and cyber operations to surveillance and fraud,” a description that matches the seven-category framework Anthropic uses internally. Reuters also reported that Anthropic disrupted a suspected Russia-linked cyber espionage campaign against Ukraine and a separate attempt to use Claude models to research biological weapons, as covered by Livemint’s report on the wire story.
Anthropic tracks individual threat clusters under internal case IDs it calls Generative Threat Groups, or GTGs. The September report names at least ten of them, spanning Russian state espionage, Chinese exploit development, election interference in Malaysia and financially motivated fraud rings. That level of individual case tracking is new compared with Anthropic’s earlier disclosures, and it is the main reason this report runs longer and covers more ground than anything the company has published before.
Seven Harm Categories, Eight Months of Activity
The scale of individual cases is where the report gets specific. One Russian-linked cluster, tracked as GTG-20006 and associated with the espionage group known as Midnight Blizzard, targeted more than 20 distinct organizations. A separate cluster linked to the ShinyHunters extortion collective, GTG-50014, had analyzed 1.8 million Android APK files as part of a malware-development pipeline before Anthropic cut off its access.
Other clusters focused on influence rather than intrusion. GTG-54002, described as a commercial influence-as-a-service operation, ran 70 fabricated news websites and more than 250 inauthentic social media accounts that had published upward of 8,900 articles. GTG-84005 ran a Malaysian election-manipulation operation built around more than 1,000 fake accounts. A Russian state-media editorial pipeline, tracked as GTG-24015, used Claude to help produce content across state-aligned outlets. Anthropic also flagged GTG-04001, a Russian foreign information manipulation and interference (FIMI) operation aimed at the Central African Republic, extending the geographic reach of the report well beyond the US, European and Middle Eastern targets that dominate most nation-state AI coverage.
| Cluster / Actor | Category | Country / Region | Key Figure |
|---|---|---|---|
| GTG-20006 (Midnight Blizzard-linked) | Cyber espionage | Russia | 20+ organizations targeted |
| GTG-50014 (ShinyHunters-affiliated) | Cybercrime tooling | Multiple | 1.8M Android APKs analyzed |
| GTG-10007 | Exploit development | China-based | 50 organizations targeted |
| GTG-54002 | Influence-as-a-service | Multiple | 70 fake outlets, 8,900+ articles |
| GTG-84005 | Election manipulation | Malaysia | 1,000+ fake accounts |
| GTG-04001 | FIMI operation | Central African Republic | Russian state-linked |
| Yemen-based weapons cluster | Conventional weapons | Yemen | Missile and drone-related guidance |
Anthropic’s own count of exfiltrated data in the cyber category is one of the more concrete numbers in the report: one campaign pulled more than 400,000 telecom records, and a separate intrusion exfiltrated over 2,100 Azure Active Directory token sets in a 34-hour window. Those figures suggest the automation gains from AI assistance are showing up in both speed and volume, not just in the sophistication of individual attacks.
Iran: Propaganda Offices and Naval Targeting Data
Iran shows up in the report in two distinct roles. The first is influence operations. Anthropic’s report and related coverage describe an Iranian propaganda office, tied to the country’s Ministry of Culture and based in Mashhad, that used Claude to help plan and refine influence campaigns aimed at audiences outside Iran. The second role is more operational: Iran-linked actors also sought targeting recommendations tied to US naval forces, using Claude in a support capacity rather than as an autonomous weapons system.
Anthropic frames this as consistent with what it has seen from Iranian state-linked activity in prior disclosures: propaganda and information operations remain the primary use case, with occasional attempts to extract military-relevant analysis. The company says it disrupted the accounts involved once the pattern was identified, though it has not disclosed how long the accounts were active before detection.
The influence operations tied to Iran did not stay within its own borders. Anthropic’s report says influence campaigns linked to Russia, Iran, Turkey and operators across the Gulf, South Asia, Africa and Europe collectively targeted audiences on six continents, a scope that makes this one of the most geographically distributed misuse patterns Anthropic has documented in a single report.
Russia: From State Espionage to Freelance Operators
Russia accounts for the largest share of named clusters in the report. Beyond the Midnight Blizzard-linked GTG-20006 espionage cluster, Anthropic also disrupted a group it tracks as a Russian “freelance” operation, separate from state-directed clusters, that offered AI-assisted attack capabilities on a for-hire basis. That distinction matters for defenders: it means the same underlying AI misuse techniques are propagating from state intelligence services down to independent operators who sell access rather than running campaigns themselves.
Reuters reported that Anthropic disrupted a suspected Russia-linked cyber espionage campaign against Ukraine as part of this disclosure, tying the report directly to the ongoing war rather than treating it as generic nation-state activity. Anthropic also documented a Russian state-media editorial pipeline (GTG-24015) that used Claude to help draft and edit content distributed across state-aligned outlets, and a financially motivated Russian cybercrime cluster (GTG-50020) operating independently of the espionage-focused groups.
This isn’t Anthropic’s first Russia-linked disclosure. In November 2025, the company disclosed a separate Chinese state-linked cluster, GTG-1002, that it described as the first largely autonomous, AI-orchestrated cyber espionage campaign attributed to a state actor, according to Google’s own threat intelligence research tracking the same shift toward agentic misuse across the industry. The September 2026 report shows that autonomy trend continuing to spread from China-linked clusters into Russian ones over the following ten months.
China: Alibaba, DeepSeek and the Distillation Race
The distillation category is where China dominates the report. Anthropic says it detected and disrupted unauthorized distillation campaigns it attributes, with high confidence, to labs based in the People’s Republic of China. The company names Alibaba specifically in connection with an attempt to extract capability from its Opus-class models to improve Alibaba’s own Qwen model family. Anthropic’s broader distillation findings also point to DeepSeek, Moonshot AI, Xiaomi and Zhipu as firms whose campaigns covertly routed user queries through Claude and used the resulting outputs to train competing systems.
Separately, Anthropic disrupted a China-based exploit-development cluster, tracked as GTG-10007, that targeted roughly 50 organizations. That cluster sits in the cyber-operations category rather than distillation, underscoring that Chinese-linked misuse in this report spans both intellectual-property extraction and direct intrusion work, not one or the other.
The distillation findings carry a commercial edge that the cyber and influence cases don’t. If Chinese labs really are routing queries through Claude to train rival models more cheaply, that goes to the core of Anthropic’s competitive position against Alibaba, DeepSeek and other fast-moving Chinese model developers, not just its security posture.
Yemen and the Houthi Weapons Pipeline
Perhaps the most alarming individual case in the report involves operators based in northern Yemen. Anthropic’s report, as described by Reuters, found that a group in Yemen used Claude to help develop software tied to weapons design, including missile systems, armed drones and bomb-related engineering work, along with associated targeting and control systems. Anthropic assesses that the group is very likely linked to Iran-backed Houthi militants, though it stops short of stating direct command-and-control ties.
This case sits at the intersection of two of the report’s seven categories, conventional weapons development and, to a lesser degree, cyber operations, since the targeting and control software work overlaps with more conventional software engineering tasks that Claude is routinely used for by legitimate developers. Anthropic’s detection relied on pattern-matching across the specific combination of technical requests rather than any single flagged query, which is consistent with how the company has described its abuse-detection pipeline in past disclosures.
The Bioweapons Threshold Anthropic Says Claude Has Crossed
The single most consequential line in the report has nothing to do with any individual country. Anthropic states that its newer Claude models can no longer be assumed to sit below the threshold for providing meaningful assistance with biological weapons. That’s a reversal from the company’s earlier public position, in which it argued its models’ uplift on bioweapons-relevant tasks remained limited enough not to require the strictest safeguard tier.
Reuters reported that Anthropic said it had broken up attempts to use Claude models to develop biological weapons as part of the same disclosure that covered the Russia-linked espionage campaign against Ukraine. Anthropic has not disclosed how many separate biological-misuse attempts it identified during the eight-month window, and the report treats biological misuse as its own category distinct from conventional weapons development, which covers the Yemen case and related kamikaze drone and missile-navigation research.
For an industry that has spent two years arguing about when frontier models would cross this line, Anthropic effectively closing the debate about its own models raises the obvious follow-up question of whether OpenAI, Google and other labs are seeing the same pattern in their own systems and simply haven’t disclosed it with the same specificity yet.
NewsGuard and NPR Put Six Chatbots to the Test
While Anthropic’s report focuses on deliberate misuse, a separate joint test by NewsGuard and NPR looked at a related but different failure mode: how often mainstream chatbots simply repeat state propaganda when asked leading questions. The test ran six AI assistants, widely reported to include ChatGPT, Gemini, Copilot, Meta AI and Grok alongside Claude, through 30 questions built from 15 false narratives that NewsGuard attributes to Russian, Chinese and Iranian state media and influence operations.
Across all six chatbots, the test found the bots correctly debunked roughly 75% of the propaganda-based questions, meaning they still repeated or failed to correct the false narrative in the remaining quarter of cases. A NewsGuard audit focused specifically on Claude found it repeated false claims 15% of the time when prompted with pro-Kremlin narratives and gave false answers to Iranian disinformation prompts in 20% of test cases.
| Metric | Result | Source |
|---|---|---|
| Chatbots tested | 6 (incl. ChatGPT, Gemini, Copilot, Meta AI, Grok, Claude) | NewsGuard / NPR |
| Test questions | 30, built from 15 false narratives | NewsGuard / NPR |
| Narratives’ origin | Russia, China, Iran state media/influence ops | NewsGuard / NPR |
| Overall correct debunk rate | ~75% across all six chatbots | NewsGuard / NPR |
| Claude repeat rate, pro-Kremlin prompts | 15% | NewsGuard Claude audit |
| Claude false-answer rate, Iranian prompts | 20% | NewsGuard Claude audit |
Read together, the two reports describe different problems that both trace back to the same models. Anthropic’s disclosure is about people deliberately trying to misuse Claude for harm. The NewsGuard test is about the model failing on its own, without any adversarial prompting beyond a leading question. Both point to the same underlying gap: safety training that catches obvious jailbreak attempts doesn’t automatically catch subtler propaganda repetition.
How OpenAI, Google and Microsoft Compare
Anthropic isn’t the only lab publishing this kind of disclosure, though its September 2026 report is the most granular one released so far. OpenAI has partnered directly with Microsoft Threat Intelligence to disrupt state-affiliated actors misusing its models, according to Cybersecurity Dive’s coverage of the partnership. That joint effort disrupted five state-affiliated actors linked to Russia, Iran, North Korea and China, though the disclosed use cases skewed toward lower-level precursor tasks such as open-source research, translation and basic code debugging rather than the weapons-development and bioweapons cases Anthropic describes.
Google’s Threat Intelligence Group has published its own tracking of adversarial AI use, describing a shift from simple prompting toward agentic AI workflows. In one case documented in Google’s research, attackers compromised a cloud resource, then used AI agents to plan, build and execute a mass credential-harvesting campaign in under six hours, a timeline that would have taken a human team considerably longer to execute manually.
| Company | Disclosure Type | Notable Finding |
|---|---|---|
| Anthropic | Standalone threat intelligence report (Sept. 2026) | 7 harm categories, 10+ named GTG clusters, bioweapons threshold statement |
| OpenAI + Microsoft | Joint disruption disclosure | 5 state-affiliated actors (Russia, Iran, North Korea, China) disrupted |
| Google (GTIG) | Ongoing threat tracker | Agent-enabled credential-harvesting campaign executed in under 6 hours |
| Meta / xAI | No comparable public disclosure found as of Sept. 2026 | — |
The pattern across all three labs that do publish is the same: agentic, multi-step misuse is replacing single-prompt attacks, and each company is disclosing it on its own schedule with its own naming conventions, which makes cross-company comparison harder than it should be for defenders trying to track a threat actor across platforms.
From “Vibe Hacking” in 2025 to Autonomous Espionage in 2026
Anthropic’s first major misuse disclosure, published on August 27, 2025, described a very different threat landscape. That report detailed a cybercriminal extortion operation nicknamed “vibe hacking” that used Claude Code to automate attacks against at least 17 organizations across healthcare, emergency services, government and religious institutions, with ransom demands sometimes exceeding $500,000. The same report described North Korean operatives using Claude to fabricate professional identities and pass technical interviews at Fortune 500 technology companies, and a lone cybercriminal selling AI-generated ransomware variants on dark web forums for $400 to $1,200 per package.
Thirteen months later, the scale and sophistication documented in the September 2026 report are categorically different. Where the 2025 report centered on individual criminals and small extortion crews, the 2026 report names state-linked clusters running for months, coordinating influence campaigns across six continents and, in the Yemen case, contributing to physical weapons development. The distance between those two reports is a rough proxy for how fast frontier-model misuse has scaled in a little over a year.
Market and Policy Stakes for Anthropic
Anthropic has built part of its brand identity around being the safety-focused alternative to OpenAI and Google, and this report cuts both ways for that positioning. On one hand, publishing granular, named-cluster detail on state-linked misuse, including a direct admission that its models have crossed the bioweapons-assistance threshold, is the kind of transparency regulators and enterprise customers say they want. On the other hand, it hands critics a concrete list of failures rather than a vague acknowledgment of risk.
The distillation findings involving Alibaba, DeepSeek, Moonshot AI, Xiaomi and Zhipu also carry direct commercial weight. If those firms really are extracting capability from Claude’s Opus-class models to train Qwen and competing systems more cheaply, that undercuts the R&D-cost advantage Anthropic is trying to protect against faster-moving, lower-cost Chinese competitors. Expect the distillation section of this report to feature in Anthropic’s ongoing conversations with US policymakers about export controls and API access restrictions for foreign labs.
What Security Researchers and Newsrooms Are Saying
Anthropic framed the scope of the disclosure directly in its own announcement. In a post on X, the company’s official account said: “We’re publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them,” according to Anthropic’s own announcement.
Anthropic’s Threat Intelligence team described the underlying detection effort on its dedicated landing page: “Over the past eight months, our Threat Intelligence team identified and disrupted operations in which threat actors tried to use Claude for malicious activity,” the team wrote, in comments hosted on Anthropic’s Threat Intelligence page.
Reuters, whose reporters reviewed the report ahead of publication, wrote that “Anthropic said in a threat intelligence report on Thursday that several actors had used its Claude AI models for activities ranging from weapons development and cyber operations to surveillance and fraud,” a framing echoed across the wire coverage that followed, including Livemint’s republication of the Reuters report. Reuters separately reported that “Anthropic broke up attempts to use its Claude models to develop biological weapons and carry out a suspected Russia-linked cyber espionage campaign against Ukraine, the AI heavyweight said in a report published on Thursday.”
Five Predictions for AI Misuse Through 2027
First, expect more labs to start naming individual threat clusters the way Anthropic does with its GTG system. Once one major lab sets that transparency bar, enterprise customers and regulators will start asking OpenAI, Google and Microsoft why their own disclosures remain comparatively vague.
Second, distillation disputes between US and Chinese labs will escalate from technical accusations into policy fights. Anthropic naming Alibaba specifically, rather than referring to “unnamed Chinese labs,” raises the odds this becomes a talking point in upcoming export-control and AI-safety legislation debates in Washington.
Third, bioweapons-threshold statements will become a standard disclosure item across frontier labs, not an Anthropic-only admission. Once one lab states its models have crossed that line, competitors face pressure to clarify their own position rather than stay silent and risk looking evasive by comparison.
Fourth, propaganda-repetition rates like the 15% and 20% figures NewsGuard measured for Claude will likely improve over the next two to three model releases, since labs treat published benchmark failures as direct engineering targets once they become public and citable.
Fifth, expect at least one more Yemen-style case, where a non-state armed group uses commercial AI for conventional weapons development, to surface publicly within the next year. Anthropic’s detection method relied on pattern-matching across combined technical requests, a technique other labs are likely already applying retroactively to their own historical query logs.
Frequently Asked Questions
What did Anthropic’s September 2026 threat intelligence report find?
Anthropic disclosed that state-linked actors from Iran, Russia and China, plus a Yemen-based group tied to Houthi militants, misused Claude models for cyber operations, influence campaigns, surveillance, weapons development and unauthorized model distillation between December 2025 and August 2026.
Did Anthropic say Claude helped build biological weapons?
Anthropic said its newer Claude models can no longer be assumed to sit below the threshold for meaningful biological weapons assistance, and that it disrupted attempts to use Claude models for biological weapons research, according to Reuters’ review of the report.
Which countries were named in the report?
The report names activity tied to Russia, Iran, China and Yemen directly, plus influence operations that also touched Turkey, the Gulf, South Asia, Africa, Europe, Ukraine, Malaysia and the Central African Republic.
What is a Generative Threat Group (GTG)?
A Generative Threat Group is Anthropic’s internal naming system for a specific misuse cluster it tracks and disrupts, similar to how cybersecurity vendors name malware families or APT groups. The September 2026 report names at least ten distinct GTG clusters.
Is Alibaba’s Qwen model connected to the Claude misuse report?
Yes. Anthropic said it detected an unauthorized distillation campaign it attributes with high confidence to a China-based lab, naming Alibaba in connection with an attempt to extract capability from Claude’s Opus-class models to improve its Qwen model family.
How did other AI chatbots perform against Russian, Chinese and Iranian propaganda?
A NewsGuard and NPR test of six chatbots, including ChatGPT, Gemini, Copilot, Meta AI, Grok and Claude, found the group correctly debunked about 75% of propaganda-based test questions built from 15 known false narratives.
Do OpenAI and Google publish similar misuse reports?
Yes, though with less granular detail than Anthropic’s September 2026 report. OpenAI has partnered with Microsoft Threat Intelligence to disrupt state-affiliated actors, and Google’s Threat Intelligence Group tracks the shift from basic prompting to agentic AI misuse.
How does this report compare to Anthropic’s 2025 disclosure?
Anthropic’s first major misuse report, published in August 2025, focused on individual cybercriminal extortion and North Korean remote-worker fraud schemes. The September 2026 report covers state-linked clusters operating across multiple countries and includes the company’s first public statement that its models have crossed the bioweapons-assistance threshold.



