ARC-AGI-3 Extracts two rankings: naked models account for less than 1%, while with the addition of the Agent framework, it reaches 36% on the first day.
2026-08-06 15:40:44
According to CoinMeta, the paper ARC-AGI-3 divides the evaluation into two separate rankings. The score of the bare model is less than 1%, but with the addition of the Agent framework, the score on the first day reaches 36%. The official ranking prohibits external harness; all models use the same minimal prompt words, and no tools are provided. What is measured is the model's native intelligence without any external support. In contrast, the community ranking allows for harness; scores are reported by the participants themselves, and ARC Prize does not perform any default verification. The paper explicitly warns that "scores from the community ranking should not be interpreted as evidence of AGI progress." On the official ranking, the scores of cutting-edge models are all below 1%.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.