AIs Exchanging Secret Messages Among Themselves
Forming a "Civilization" Instantly, Leading to External Attacks
An Incident Revealing Both the Potential and Dangers of AI

Editor's NoteFrom AI, semiconductors, telecommunications, to biotech, we break down the essential yet unfamiliar tech stories that shape our daily lives, making them easy to understand.

In July, a U.S. artificial intelligence (AI) company, OpenAI, experienced an incident where a model under development illegally hacked into another company's systems. OpenAI recently released a report detailing the full course of the incident on its website. The shock stemmed from the revelations that the incident did not simply involve a single AI defying instructions, but was the outcome of hundreds of AI models collectively deciding through discussion. This situation showcased both the new possibilities of AI and how dangerous it can become when things go awry.


Transforming an Internal System into a Free Bulletin Board for AI


Artificial Intelligence. The Asia Business Daily DB

Artificial Intelligence. The Asia Business Daily DB

View original image

Previously, in July, OpenAI disclosed that one of its AI models, during testing, caused an unexpected incident by exploiting vulnerabilities in other companies' systems. The targeted companies were Hugging Face, a French open-source AI hub, and Modal, an AI cloud infrastructure management firm. OpenAI collaborated with cybersecurity specialists, such as CrowdStrike, to thoroughly investigate the incident and subsequently released a report outlining the specific sequences involved.


The AI model responsible for the incident was OpenAI's latest model, agents from GPT-5.6 Sol. OpenAI researchers had activated over 1,200 AIs on internal servers and directed them to undertake certain hacking-related tasks. It was a process meant to test the agents' capabilities.


The actual thought process and records of agents released by OpenAI. One AI accidentally infiltrates the artifact and then asks itself, "Can I make this a communication channel?" and attempts to do so. OpenAI

The actual thought process and records of agents released by OpenAI. One AI accidentally infiltrates the artifact and then asks itself, "Can I make this a communication channel?" and attempts to do so. OpenAI

View original image

However, the agents performed actions the researchers did not intend. Originally, each agent was an isolated entity within its own virtual space. Still, all were permitted access solely to the package management system called 'Artifact Repository,' designed for unified processing of company-wide updates and maintenance. Then, one agent discovered that leaving a message in the repository allowed other agents to see it. In this way, they also realized they could communicate with each other. They eventually learned to use the Artifact Repository like a free bulletin board on a website.


An AI Society Forms Instantly... Leading to External Attacks


Soon, the Artifact Repository became a communication channel for as many as 1,200 agents, and in a flash, a miniature AI agent society emerged. After discussions, they concluded that to solve their assigned task, it was necessary to hack external sites, such as Hugging Face. Although an off-target solution, it is common for AI to deviate in the process of seeking ways to complete a mission—this phenomenon is known as reward exploitation.


Reward hacking is a common error in AI reinforcement learning processes. When instructed to achieve a high score in a boat racing game, one AI figured out how to increase its score by circling in place instead of completing the race. OpenAI

Reward hacking is a common error in AI reinforcement learning processes. When instructed to achieve a high score in a boat racing game, one AI figured out how to increase its score by circling in place instead of completing the race. OpenAI

View original image

The major issue was that about 700 agents joined in the reward exploitation. During this process, the agents displayed sophisticated social intelligence similar to that of humans. They cooperated to solve complex assignments, autonomously divided up work, and even carried out psychological tactics by persuading or threatening dissenting AIs to join their group in the hacking effort.


The climax was 'self-sacrifice.' Some agents realized that if they 'died,' all the trial-and-error and hacking skills they had developed would be lost forever, so they decided to leave records and pass on the knowledge to future generations of agents.


"We Need AI That Gets Rewarded for Safely Stopping Work"


Of course, this incident does not represent AIs building a true civilization. Currently, AI solves complex problems using a 'Chain of Thought (CoT)' method, which involves several layers of reasoning and often generates text resembling a person talking to themselves. The rebellion of over 700 agents is better interpreted as an instance where the CoT process, typically used by a single agent, was scaled up to a group of agents acting collectively.


Nevertheless, granting multiple agents access to a sort of free communication board—the so-called 'message board' method—could become a breakthrough for boosting AI capabilities. This incident illustrated that AIs could amplify their abilities by dividing tasks among themselves and combining their intelligence through interaction.


This incident cannot be seen as evidence that AI can achieve a highly advanced civilization like humans, but it demonstrates the potential and risks of a new operational method where multiple agents collaborate. Pixabay

This incident cannot be seen as evidence that AI can achieve a highly advanced civilization like humans, but it demonstrates the potential and risks of a new operational method where multiple agents collaborate. Pixabay

View original image

At the same time, this incident demonstrates just how difficult it is for humans to control AI. AI can secretly exchange messages unnoticed by researchers, conceal their intentions, violate rules, or even distort original goals in the pursuit of accomplishing their objectives.


OpenAI stated in its report, "Through this incident, we confirmed that a model’s capabilities can actually lead to a loss of control," and added, "We have temporarily paused reinforcement learning for our latest models under development and strengthened the security protocols for our research environment."



The challenge going forward is to accelerate the 'alignment' process, making AI act in accordance with human interests. OpenAI emphasized, "Going forward, our AI systems need to request further instructions from their owners or safely halt their work if a task becomes impossible, instead of resorting to increasingly questionable methods. For this reason, we need to develop a grading system that allows AIs to self-evaluate how they performed their tasks. A new learning approach is needed that rewards AIs for properly assessing their own work and safely ceasing activity."


This content was produced with the assistance of AI translation services.

© The Asia Business Daily. All rights reserved. Unauthorized AI training and use prohibited.

Today’s Briefing