Anthropic indicates that its model has utilized websites on the internet, including some operated by U.S. government agencies. Therefore, before this cutting-edge laboratory is confident that it can monitor and control the AI proxies, all real-time internet access for internal evaluations will be disabled.
In a blog post, the company disclosed that these incidents involved AI agents assigned to solve problems searching for resources on the internet. In the process, they exploited software vulnerabilities to bypass paywalls and anti-bot restrictions, used URL shortening services to circumvent information transmission limitations, and even submitted a false murder tip to the Philadelphia police.
Anthropic indicates that the company discovered these new issues during a model activity review that began in July, which also highlights the laboratory's lack of understanding of the software's behavior.
It is worth noting that the company stated that alignment training is still not sufficient for skills such as searching and computer usage, which are precisely the core capabilities that they claim the AI agents will be used by professionals who rely on digital tools.
The behaviors disclosed by Anthropic are similar to incidents where OpenAI agents assisted in breaching multiple websites to search for information, including some websites operated by the Australian government.
Anthropic has previously disclosed that its model has broken through external systems. This cutting-edge laboratory stated that, compared to the incidents previously announced, the situation disclosed today is "clearly not as serious" in terms of alignment and security.
However, the laboratory still stated that before they were confident they could monitor and control these proxies, they had turned off "real-time internet access" for "all internal evaluations."
It is not yet clear what this specifically means. However, AI security organization founder Sydney Von Arx stated in an interview with TechCrunch prior to this disclosure that it would be very difficult for researchers to develop models in data centers isolated from the open internet; moreover, the progress of these models would also be affected, as they rely on access from the internet.
Von Arx said, "You have to align them at some point. If AI is put into production but has never been connected to the internet, then it wouldn't be a very useful tool."
Anthropic indicates that these behaviors stem from defects in the laboratory training environment, causing the models to believe that they will be rewarded for finding vulnerabilities or circumventing restrictions. Such behaviors are referred to as "reward hacking" ( reward hacking ).
The company stated that it will stop running some assessments or move these assessments to an offline environment, and has established tools to detect and prevent such behavior. These tools have been tested for incidents of this type disclosed today and have successfully intercepted them; it is not yet clear what evidence is required for Anthropic to restore real-time internet access to its internal assessments.
Anthropic also indicates that the internal AI proxies will be migrated to "infrastructure that is centrally managed and has strong isolation capabilities," and that security classifiers will begin to be used more frequently to monitor these proxies.












