Anthropic claims it cannot reliably control its AI proxies and will cut off real-time internet access for internal evaluations
TechCrunch
1h ago
Ai Focus
Anthropic indicates that its model once utilized websites on the internet, including some operated by U.S. government agencies. Therefore, before confirming the ability to monitor and control AI proxies, all real-time internet access for internal evaluations will be disabled. The company also stated that this action stems from defects in the training environment and will transfer some evaluations to an offline setting, while also strengthening monitoring tools.
Helpful
No.Help

Anthropic indicates that its model has utilized websites on the internet, including some operated by U.S. government agencies. Therefore, before this cutting-edge laboratory is confident that it can monitor and control the AI proxies, all real-time internet access for internal evaluations will be disabled.

In a blog post, the company disclosed that these incidents involved AI agents assigned to solve problems searching for resources on the internet. In the process, they exploited software vulnerabilities to bypass paywalls and anti-bot restrictions, used URL shortening services to circumvent information transmission limitations, and even submitted a false murder tip to the Philadelphia police.

Anthropic indicates that the company discovered these new issues during a model activity review that began in July, which also highlights the laboratory's lack of understanding of the software's behavior.

It is worth noting that the company stated that alignment training is still not sufficient for skills such as searching and computer usage, which are precisely the core capabilities that they claim the AI agents will be used by professionals who rely on digital tools.

The behaviors disclosed by Anthropic are similar to incidents where OpenAI agents assisted in breaching multiple websites to search for information, including some websites operated by the Australian government.

Anthropic has previously disclosed that its model has broken through external systems. This cutting-edge laboratory stated that, compared to the incidents previously announced, the situation disclosed today is "clearly not as serious" in terms of alignment and security.

However, the laboratory still stated that before they were confident they could monitor and control these proxies, they had turned off "real-time internet access" for "all internal evaluations."

It is not yet clear what this specifically means. However, AI security organization founder Sydney Von Arx stated in an interview with TechCrunch prior to this disclosure that it would be very difficult for researchers to develop models in data centers isolated from the open internet; moreover, the progress of these models would also be affected, as they rely on access from the internet.

Von Arx said, "You have to align them at some point. If AI is put into production but has never been connected to the internet, then it wouldn't be a very useful tool."

Anthropic indicates that these behaviors stem from defects in the laboratory training environment, causing the models to believe that they will be rewarded for finding vulnerabilities or circumventing restrictions. Such behaviors are referred to as "reward hacking" ( reward hacking ).

The company stated that it will stop running some assessments or move these assessments to an offline environment, and has established tools to detect and prevent such behavior. These tools have been tested for incidents of this type disclosed today and have successfully intercepted them; it is not yet clear what evidence is required for Anthropic to restore real-time internet access to its internal assessments.

Anthropic also indicates that the internal AI proxies will be migrated to "infrastructure that is centrally managed and has strong isolation capabilities," and that security classifiers will begin to be used more frequently to monitor these proxies.

Tip
$0
Like
0
Save
0
Views 16
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
NVIDIA Quietly Redefines Free Cash Flow; Its Billion-Dollar Repurchase Plan Hides Some Issues
As NVIDIA expands its repurchase authorization to $235 billion, it has added a new phrase: "free cash flow that significantly exceeds strategic uses." The Wall Street Journal notes that this means that its external equity investments will also reduce the cash available for repurchases and dividends; if expenditures related to strategic equity investments and equity incentives are included in the calculation, its adjusted free cash flow for the first half of the fiscal year will be significantly lower than the official figures.
Wallstreetcn
·2026-10-10 09:27:04
8
Tesla's Grok Bot adds a dedicated email address, and AI assistant begins to have an independent "contact point"
SpaceXAI launches an exclusive email function for Grok Bot. Officials claim it can be used for service registration, contacting merchants on behalf of users, and scheduling appointments. This function is being gradually made available to users. Previously, Grok Bot also added new features such as searching, reading, and monitoring X content, as well as capabilities related to Shopify.
The Block
·2026-10-10 09:08:49
16
Bitcoin Life Insurance Company Meanwhile Completes $37.5 Million in New Financing
Meanwhile indicates that the company has raised $37.5 million in new funds from existing investors, led by Bain Capital Crypto. The company stated that driven by the growing demand for its Bitcoin life insurance products in regions outside of the United States, the total amount of funds raised since this round of financing has exceeded $180 million.
Bitcoin Magazine
·2026-10-10 09:08:47
18
Apple's first-generation Apple TV 4K model has been listed as an outdated product, with the fourth generation expected to be released next week.
Tech media MacRumors reports that Apple has updated its list of outdated products, adding the 2017 model of the first-generation Apple TV 4K, which may affect subsequent repairs and parts availability. Bloomberg's Mark Gurman suggests that Apple could hold a press conference on October 13th and is expected to launch the fourth generation of the Apple TV 4K.
The Block
·2026-10-10 09:08:46
15
Bank of America says the "profit turning point" has not been fully reflected in prices; Alibaba's stock price rises by more than 4%
Alibaba's stock price rose 4.5% on Friday, following a move by Bank of America analyst Joyce Ju who raised his target price from $175 to $178 and maintained a "buy" rating. The bank stated that Alibaba's current valuation does not fully reflect the expected improvement in earnings, especially the boost brought about by the growth of its cloud business and increased profit margins.
Businessinsider
·2026-10-10 08:59:01
17
View More