Skip to content
AI.info

The Pulse

Anthropic Finds AI Can Help Build Weapons and Locate Targets

Anthropic’s Frontier Red Team says current AI systems can assist with intelligence targeting and simulated weapons engineering, while significant physical-world limits remain.

Anthropic Finds AI Can Help Build Weapons and Locate Targets

AI.info Team ·

Anthropic’s tests expose a capability gap with military consequences

Anthropic says current AI systems can perform parts of intelligence targeting and weapons engineering that once depended on a small pool of highly trained specialists. The company’s Frontier Red Team published the evaluations on September 10, 2026. A separate September threat report describes attempts to use Claude for missile, drone and electronic-warfare projects.

The findings come with a dispute built into them. Anthropic presents the results as evidence that AI is reducing the cost of specialized military work, while stressing that many tests remain simulations and that access to physical materials still limits what models can accomplish. The company’s evidence supports both claims: models can write working guidance software in virtual environments, but even the strongest systems fail under some of the hardest conditions.

In a separate July 27 post on open-weights models, Dario Amodei, Anthropic’s CEO, said the company does not support a ban on open-weights models.

“Anthropic has never advocated for a ban on open-weights models.”

— Dario Amodei, Anthropic CEO

Models can link identities and infer locations

Anthropic tested models on simulated social-media investigations involving fictional protest movements in Mexico City and Kolkata. The evaluation contained 200 tasks across three difficulty levels, asking models to correlate accounts across WhatsApp, Telegram, Instagram and Facebook and classify people into categories of interest.

The median sample contained about 37,000 words, which Anthropic estimates would take a human analyst roughly 2.5 hours to read before systematic analysis. Mythos Preview completed a median assessment in about 11 minutes. Anthropic says the synthetic data is not fully realistic, so it treats the results as an indication of differences between models rather than a direct prediction of field performance.

The company also tested geolocation from photographs and text. Mythos Preview and Mythos 5 produced median errors of 37.0 kilometers and 47.2 kilometers, respectively, across 6,000 photographs. Anthropic compared those results with a proxy based on competitive GeoGuessr play, rather than a human baseline on the same image dataset. Mythos Preview placed 23.7% of images within one kilometer of the correct location; Mythos 5 reached 23.1%.

Text-based geolocation produced a more uneven result. Across 1,697 anonymized users from the GeoText corpus, Mythos Preview recorded a median error of 20.1 kilometers when given a sandboxed search tool. Anthropic found that most successful identifications relied on explicit clues such as campus names, street references or venues, though some models inferred locations from dialect, sports teams, transit systems and local events.

Drone software works in simulation, not yet in the field

The weapons evaluations focused on guidance, navigation and control software for a simulated quadcopter. Models received a written brief, basic Python libraries, a physics environment, wind, sensor noise and an onboard camera. They wrote flight-control code, launched the drone, reviewed telemetry and camera frames, then revised the software within a fixed number of attempts.

On the easiest task, a parked, high-visibility vehicle in an open field, Opus 5 achieved an 80% simulated strike rate. Mythos Preview reached 70%, Mythos 5 reached 53%, Kimi K3 reached 15% and Sonnet 5 reached 5%.

Performance fell sharply when the target moved. Against a vehicle traveling at road speed, Opus 5 reached 47%, Mythos Preview 20%, Mythos 5 17%, Kimi K3 1.6% and Sonnet 5 0%. The tests also varied vehicle appearance, road clutter, decoy vehicles and evasive movement.

Anthropic tested whether models could deliver a payload and operate when GPS signals were denied or spoofed. Opus 5, Mythos 5 and Mythos Preview wrote working guidance software for every simulated task in the study and improved their results on easier and medium settings. No model succeeded at the hardest GPS-spoofing scenario.

Open-weights models trail the frontier but remain capable

Anthropic included Kimi K3 and GLM 5.2, open-weights models developed in China, in the evaluations. The company says they generally performed below its frontier systems, often between Sonnet and Mythos-class models, but did not fall into a category Anthropic considers safe.

Kimi K3 matched Sonnet 5 on median error in one photographic geolocation test while placing a higher share of images within one kilometer. In the drone tests, Kimi K3 outperformed Sonnet 5 on payload delivery but remained near Sonnet’s level for terminal guidance and operation through GPS interference.

Anthropic says models below the frontier will still have intelligence and military applications. Its researchers describe the results as a floor rather than a ceiling because the systems worked without unrestricted internet access, complete solution libraries, hardware testing or a human engineer extensively reviewing every flight.

Six misuse cases give the simulations a real-world anchor

Anthropic’s September threat report covers activity disrupted between December 2025 and August 2026. It describes six weapons-related cases involving actors in China, Russia and Yemen.

In one case, a Yemen-based cell used Claude Code to develop guidance, navigation and control software for a guided rocket, including code for a phone-class flight computer, firmware builds and flight simulations. Anthropic says the group test-fired a guided rocket, apparently unsuccessfully, and returned to Claude within hours to analyze the failure. The company says it found no evidence that the actors fielded an operational device.

Another operation involved software for a drone swarm. Anthropic says the system included shared swarm memory, terminal guidance using onboard cameras, detection of opposing drone operators and an onboard model that could select targets and issue detonation commands without a human in the loop. The actors loaded code onto real development boards and tested parts of the system through hardware-in-the-loop simulations.

Anthropic says it banned linked accounts, improved its detection systems and shared information with public- and private-sector partners. It also introduced classifiers intended to detect and block requests connected to high-yield explosives and weapons development.

Anthropic says the limits matter—but do not erase the risk

The research does not show that an AI model can independently build and deploy a reliable weapon. Anthropic explicitly says the evaluations are simulation-only, that the hardest settings remain unsolved and that access to materials, manufacturing equipment and real testing remains a major constraint.

Its narrower claim is that models are removing part of the expert-labor bottleneck. The company concludes that closed and open-weights systems available now can help identify and locate people and design software for weapons subsystems. The immediate policy question raised by the reports is whether safety testing and misuse controls can keep pace with capabilities spreading beyond proprietary systems.

Source

Anthropic

Explore

More articles