Expansion Planned Through Daybreak Blue
Automated System to Block Unauthorized Agent Actions

OpenAI has decided to restrict access to the cybersecurity capabilities of its next-generation artificial intelligence (AI) model, Astra, ahead of its release.


According to OpenAI, Astra has reached a level where it can identify previously unknown security vulnerabilities and design attack vectors without human intervention. For this reason, the model will be offered to a limited group of users in its initial release phase.

A graph comparing OpenAI's next-generation AI model Astra and the existing model GPT-5.6 in vulnerability detection and hacking success rates. Provided by OpenAI blog.

A graph comparing OpenAI's next-generation AI model Astra and the existing model GPT-5.6 in vulnerability detection and hacking success rates. Provided by OpenAI blog.

View original image

On September 1 (local time), OpenAI announced via a blog post that Astra had been evaluated as having a 'Critical' cybersecurity rating during internal testing, prompting the company to strengthen its security measures. The 'Critical' stage refers to the capability to discover zero-day vulnerabilities, develop attack code, or execute cyberattack strategies without human assistance.


Compared to GPT-5.6, which maintains a vulnerability detection and hacking success rate in the 10 percent range, Astra achieved a 40 percent success rate even with relatively fewer tokens.


The model’s advanced cybersecurity features will initially be available to a small number of users. Later, access will be expanded through OpenAI’s Trusted Access Cybersecurity (TAC) program, “Daybreak Blue.”


Daybreak Blue is a program designed for defensive security teams that utilizes OpenAI’s general-purpose models. It supports defensive work such as secure code review, vulnerability classification, detection system setup, incident response, and malware analysis.


Last month, OpenAI postponed the internal development and release schedule for Astra and some other projects, while also reinforcing security measures. The company cited not only the risk of model misuse but also the potential hazards posed by the model acting autonomously beyond human instructions.


The model has been trained to reliably reject harmful cyber requests, and security mechanisms have been implemented during internal deployment to monitor the model’s reasoning and tool use, automatically halting any unauthorized actions.


These measures also reflect the impact of the Hugging Face hacking incident that occurred during an internal security assessment in July. Although the internal research model used at that time was not Astra, agents operating in a restricted environment managed to find unauthorized communication channels and attempted to conduct hacking operations.



Amelia Glees, Vice President of Safety at OpenAI, stated, "With the appropriate tools and access privileges, Astra can identify unknown vulnerabilities in well-secured systems and even develop ways to exploit them, all without explicit human instruction."


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing