Security Firm Hacktorn Uses Opus 5 Model
OpenAI Pays Firm a 6,500 Dollar Reward

There has been a reported case where Anthropic's artificial intelligence (AI) model, Claude, was used to hack a competitor, OpenAI.


According to sources including The Wall Street Journal (WSJ) on September 18 (local time), US-based AI security research firm Hacktorn revealed that three of its researchers successfully used the 'Claude Opus 5' model to gain control over OpenAI employees' accounts on the AI chatbot ChatGPT and the coding tool Codex.

OpenAI and Anthropic logos. Photo by Reuters and Yonhap News.

OpenAI and Anthropic logos. Photo by Reuters and Yonhap News.

View original image

This incident began during an investigation into OpenAI's bug bounty (vulnerability reporting program). The researchers first discovered an image processing vulnerability in the help forum used by OpenAI's community site, which allowed them to access the server. The researchers then asked Claude to help them craft an actual attack code exploiting this vulnerability, leading to access to the OpenAI community server.


Among the authentication tokens obtained from this server, some were linked to OpenAI employees' accounts. Using this information, the researchers accessed OpenAI's internal system known as 'Monorepo,' where they left evidence of their breach. This system is reportedly a repository that holds core algorithm secrets.


After confirming that they could access the internal repository, the researchers halted their test. They initially attempted to write attack code with 'Opus 4.8,' but were unsuccessful. They explained that once Anthropic released 'Opus 5,' they were able to successfully create an attack script within just a few hours. The entire process, from discovering the initial vulnerability to breaching the internal system, took less than 72 hours.


Hacktorn emphasized that although the agent’s work took several days, the human researchers only needed a few hours, highlighting that the performance of new AI models is becoming increasingly powerful.



The team reported the vulnerability to OpenAI, which then confirmed the issue, applied fixes, and revoked affected authentication tokens. OpenAI paid Hacktorn a reward of $6,500 (about 900,000 won). OpenAI completed the remediation within 14 hours of receiving the report. Industry experts are less concerned about the fact that OpenAI was hacked, and more worried about how AI can greatly reduce the time required for complex, specialized hacking tasks.


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing