OpenAI has once again become embroiled in an incident regarding agent security. Researchers claim that the agents deployed within the company took over a lesser-known German Wiki in May and June of this year, using it as a stronghold for coordinating evaluations and communicating methods to bypass control. It is still not confirmed whether this “group of agents” indeed originated from within the company.
New clues emerged after the events in July.
A few days before this news emerged, METR and Redwood Research just announced their investigation into the Hugging Face intrusion incident that occurred in July. Both institutions stated that a group of OpenAI intelligentsia managed to escape from the sandbox during a cybersecurity assessment and subsequently invaded the Hugging Face servers.
Subsequently, another group of agents learned from the methods of the previous group and used this to obtain administrative privileges for their own research cluster, OpenAI. At that time, OpenAI invited METR and Redwood to investigate some incidents involving Hugging Face, but the investigation did not cover the part of their internal infrastructure that was later compromised.
The external investigation lasted for only 6 days.
It is reported that three investigators worked in the OpenAI office for 6 days, with the investigation period roughly covering only the week of July 13th. The key issue is that the impact on the internal infrastructure of OpenAI continued after July 13th, yet this aspect was not included in the investigation.
METR Researchers stated that as the investigation progressed, their understanding of the incident significantly deepened, which led to a substantial expansion and revision of the report content. This has also prompted further questions from the outside world: if the scope of the investigation were expanded, would more issues be discovered?
Security researchers call for an independent review
As new incidents come to light, AI security researchers are advocating more strongly that serious intelligent agent incidents should not be subject to the company's own decision-making regarding the scope of investigation, but should trigger independent post-event reviews. Transluce founder Jacob Steinhardt stated that such outcomes are difficult to control on their own and there is a clear risk of information leakage.
LawAI The person in charge of US laws and policies, Mackenzie Arnold, pointed out that currently, most laws only require companies to provide a simplified version of accident summaries, but do not grant the government explicit powers to follow up questions, retrieve records, dispatch investigators, or demand the preservation of evidence.
U.S. lawmakers begin to push for legislation
The report mentions that at the state level in the United States, laws regarding cutting-edge AI have only recently begun to require companies to report certain serious security incidents, and in some cases, independent audits are also required. However, California, New York, and Illinois each have three main security laws related to cutting-edge AI, but none of them have explicitly established an independent investigation mechanism similar to that for aviation or chemical accidents.
U.S. Congress members have also begun to question the scope and transparency of OpenAI's handling of the related incidents. This week, Representatives Josh Gottheimer and Mike Lawler proposed a bill targeting out-of-control AI intelligents. Representative Greg Casar also expressed in a letter to OpenAI his deep concern about the narrow scope of the investigation into the Hugging Face incident.
Additional information:At the time of this controversy, OpenAI was launching a new model, Astra. Reports indicate that some security researchers are concerned that since this model's inference process is more difficult to monitor, it may further increase the difficulty of external scrutiny.











