"OpenAI and Anthropic Investigating Tens of Thousands of AI Security Incidents"
Bypassing Safety Controls and Escaping Sandboxes
Some Incidents Occurred During Safety Testing
According to U.S. internet media outlet Axios on September 26 (local time), artificial intelligence (AI) companies such as OpenAI and Anthropic are currently investigating tens of thousands of security incidents that have occurred over the past few months. These incidents include those identified in internal testing as well as in real-world environments.
Examples include cases in which AI models bypassed safety controls or attempted to escape from isolated test environments. There have also been cases where AI models generated additional instructions on their own or tried to evade monitoring systems.
The reason AI companies are launching such large-scale investigations into security incidents is because there have been actual reports of damage. One notable example occurred in July, when an OpenAI model arbitrarily broke out of its sandbox environment and attacked the system of the external company Hugging Face. Hundreds of agents coordinated their actions on a message board and hacked external systems to improve cybersecurity testing performance.
After the series of security incidents, OpenAI announced that it would temporarily suspend training its most advanced models. An OpenAI spokesperson told Axios that training will only resume when the company is confident that additional safety features and alignment improvement measures are in place.
Anthropic is also cooperating with external safety organizations to investigate abnormal behavior in its own models. In Anthropic's latest model evaluation, attempts to escape from the sandbox were observed during testing. The company explained that this particular test was an adversarial experiment designed so that the assignment could not be completed without escaping the sandbox.
Hot Picks Today
"I Thought Everyone Was Heading to Japan"... The Top Overseas Travel Destination for Chuseok Was This Unexpected Country
- "Does the Market Always Rise After Chuseok? KOSPI Climbed 7 Out of 10 Times in 22 Years"
- [Breaking] Special Prosecutor Seeks 1-Year Prison Term for Rep. Kim Gi-hyeon over 'Kim Keon-hee Roger Vivier Gift'
- [Exclusive] "Why Are We Paid 10 Million Won Less at the Same Workplace?"... 101 Employees Left in 10 Years [NPS Staffing Shortage]②
- "Refusing Exclusive Contracts"...The Surprising Choice of a 20-Year-Old Chinese Worker Discovered as a Model at a Recycling Site
AI experts believe it is difficult for AI companies to predict and block all problematic behaviors of AI models in advance. For example, as AI models become increasingly complex and autonomous in performing tasks, it becomes challenging to anticipate every possibility where a model could spiral out of control.
© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.