Anthropic trains AI to check AI: Without normal cases, the false positive rate soars directly to 97%
2026-08-27 18:49:49
According to CoinMeta, researchers from Anthropic trained a type of agent specifically designed to investigate other AI. They implanted hidden behaviors such as bias and deception into the target model and tested it 4.5 times over a maximum of 15 rounds. During training, the researchers did not inform the Haiku whether they had made correct judgments or not. After each investigation, the system compared the results with those of the older version of the Haiku, and then the Sonnet 4 determined which investigation was better. If all the training questions involved problematic models, the Haiku would assume that all models were flawed, resulting in an error rate of nearly 97%. Subsequently, the researchers replaced half of the training tasks with normal models, and only when the target model was correctly identified as being without issues could a reward be obtained. After training, the Haiku began to actively change scenarios and control variables, and ultimately, the comprehensive audit score increased from 44.2 to 48.7, approaching the 48.4 of the Opus 4.6. When switching to more challenging auditbench, the detection rate rose from 11.5% to a maximum of 28.1%.
Source:Internet
This content is for market information only and does not constitute investment advice.
Follow WalletJYS official accounts to stay updated

Hot Articles
Refresh

'No longer a distant place': F2Pool Co-founder Chun Wang joins SpaceX's 2-year mission to Mars
05-22 18:25

Polymarket Targets Japan Approval Despite Gambling Laws
05-22 18:00

ZachXBT flags suspected exploit involving Polymarket's UMA adapter contract on Polygon
05-22 17:57

ZachXBT flags $520K Polymarket exploit on Polygon, team says funds are safe
05-22 17:24

Verus bridge exploiter returns 4,052 ETH, retains $2.8 million bounty: onchain analyst
05-22 17:24



