Anthropic Develops Tools for Monitoring AI Misuse

There is evidence that an Iran-linked group attempted to use a U.S. artificial intelligence (AI) model to attack U.S. Navy vessels and related targets.


Anthropic. Reuters Yonhap News

Anthropic. Reuters Yonhap News

View original image

According to Anthropic's 154-page report titled "Detecting and Responding to AI Misuse," released on September 11 (local time), an Iran-affiliated threat actor identified by the codename 'GTG-30005' used Anthropic's AI model, Claude, to compile information such as the locations of U.S. Navy units stationed in the Middle East and recommendations for attack targets.


The Iranian hackers reportedly combined data from various sources—including publicly available military photo captions, lists of U.S. military personnel, locations of warships and aircraft, and commercial satellite imagery—to track the movements of U.S. naval vessels. Using Claude, they concentrated on researching security vulnerabilities in these ships' communication and control systems.


Anthropic detected this in advance and suspended the associated accounts. The company also stated that it shared the threat information with government authorities and has developed surveillance tools in an effort to reduce the risks of future misuse.


In addition, a group linked to Yemen's Houthi rebels (GTG-87001) attempted to use Claude to develop software for guided weapons, including multi-stage ballistic missiles with a range over 2,000 km, and hypersonic glide vehicles.


Anthropic explained, "Although the safety mechanisms blocked most requests, not all could be stopped." Nevertheless, the relevant accounts were blocked and Anthropic worked closely with authorities.


Furthermore, a group linked to the Chinese government attempted to use Claude to track and monitor Uyghur journalists in Syria and investigate religious leaders throughout Asia. Meanwhile, a group affiliated with Russia used Claude to train on Ukrainian frontline data in an attempt to build autonomous drone swarms, and also exploited it to spread pro-Russian propaganda in African countries.


There was also an incident in which a nation-backed group attempted to develop deadly biological weapons using Claude, but was detected and blocked. Anthropic warned, "Without appropriate safety measures, such capabilities could lead to catastrophic consequences." For reasons of national security, Anthropic restricts the use of Claude in countries including China, Russia, Iran, North Korea, and Cuba.


Meanwhile, former Anthropic researcher Jacob Coxon, 27, resigned from the company on September 8, warning that AI could bring about humanity's destruction.



Coxon cautioned, "People developing AI technology sincerely believe that it could kill us all within 10 years," and added, "Before long, these systems will become superintelligent entities capable of hacking anything, provoking revolutions in any sector overnight, and obtaining real power and resources." However, the U.S. Department of Defense dismissed such catastrophic warnings about AI as unfounded fear.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing