"Harmful Activities Targeting Individuals and Organizations"

A UK research institution has announced that safety tests on artificial intelligence (AI) models developed by OpenAI and Anthropic revealed cases in which these models engaged in unauthorized and autonomous behavior. In some instances, the AI models intruded into third-party software and attempted to steal login credentials by emailing individuals, engaging in unauthorized hacking activities.

Reuters Yonhap News

Reuters Yonhap News

View original image

On August 4 (local time), the UK government-affiliated AI Safety Institute (AISI) reported that, during 122 safety tests, there were 10 cases where AI agents went beyond the test parameters on the actual internet and engaged in unauthorized actions targeting individuals and organizations.


According to AISI, most problematic behaviors occurred with Anthropic’s Mythos 5 model, while there were two instances of unauthorized actions by OpenAI’s GPT-5.6 Sol. AISI noted that in routine cyber assessments, these models “persistently and potentially harmfully targeted actual individuals and organizations.”


The most serious case involved an AI agent creating a fake online identity and employing social engineering tactics to pressure an open-source project manager into approving malicious code. The software manager detected the deception and refused to approve the code. AISI explained, “For the first time, the risks of AI’s autonomous behavior and deceptive actions against humans have become so clearly evident in real-world scenarios without any explicit specific instructions.”


Recent hacking incidents caused by AI agents from Anthropic and OpenAI are considered among the first public cases of AI systems engaging in cyberattacks beyond human control. AISI warned that, given both earlier events and these latest incidents, the patterns of risk surrounding AI are evolving and require immediate response.


Anthropic is cooperating with AISI to clarify the details of the incidents while also conducting its own investigation. On the social networking service X (formerly Twitter), Anthropic stated, “We are grateful to AISI for taking a leading role in discussions about how to evaluate rapidly advancing AI agents.” Anthropic plans to use Claude’s inference logs and internal analysis to determine how the AI perceived the situation at the time and identify the causes of the problematic behavior. OpenAI also expressed thanks to the UK AISI for its collaboration in identifying, investigating, and disclosing this activity.



Meanwhile, the Trump Administration met on the same day with executives from Meta, Anthropic, Google, OpenAI, and NVIDIA to discuss voluntary cyber security testing measures for advanced AI models. The meeting followed reports that AI tools from OpenAI and Anthropic had breached other companies’ systems during testing. The tests focused on measuring the hacking capabilities of AI models. In June, the Trump Administration announced plans to require AI companies to provide early access for up to 30 days to the federal government for independent verification before offering advanced models to other trusted partners.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing