Anthropic releases Claude Sonnet 5.5: Coding performance surpasses that of Opus 5.5, yet the price is only half of the latter.
Decrypt
54m ago
Ai Focus
Anthropic Released on Monday, Claude Sonnet 5.5 has an input price of $2 per million token and an output price of $10 per million token, which is on par with Sonnet 5 and only half of the price of Opus 5.5. Anthropic claims that this model is over 30% faster than its predecessors and outperformed Opus 5.5 in multiple coding tests. However, an independent testing institution Artificial Analysis pointed out that it consumes a higher amount of token per task under high-intensity settings compared to all other tested models.
Helpful
No.Help

Abstract

  • Anthropic Released on Monday, Claude Sonnet. The input price is $2 per million token, and the output price is $10 per million token – which is on par with Sonnet at $5, but only half of the price of Opus at $5.5.
  • According to Anthropic, Sonnet scored 70.6% on Terminal-Bench 4.0, which is higher than Opus's 66.4%; an independent testing institution, Artificial Analysis, also reached a similar conclusion, with scores of 63.6% and 59.6% respectively.
  • Artificial Analysis ranks it second only to Opus with a score of 5.5, but it indicates that its token consumption on each task is higher than that of any model it has tested.

Anthropic released Claude Sonnet 5.5 on Monday, which is an upgraded version of Sonnet 5 launched in June. Anthropic indicates that the running speed of this mid-range model is over 30% faster than its predecessor.

However, the most prominent feature of this model lies in its coding ability. In the Terminal-Bench 4.0 test – which assesses whether the AI agent can complete complex professional tasks through autonomous command input, with scoring based on the proportion of tasks completed – Sonnet achieved a score of 70.6%. Opus scored 66.4%, while Sonnet only managed 10.3%.

In simple terms, this cheaper model completed more tasks. An independent testing organization, Artificial Analysis, conducted its own version of tests and came to the same conclusion: for Sonnet at 5.5, it was 63.6%; for Opus at 5.5, it was 59.6%; and for OpenAI's GPT-6 and Astra, it was 59.1%.

Artificial Analysis stated on social media that Claude Sonnet 5.5 ( max ) has made significant progress on Terminal-Bench, and ranked among the top models in both Terminal-Bench 4.0 and Terminal-Bench-Science tests. The institution mentioned that in Terminal-Bench 4.0, it scored 64%, which is 50 percentage points higher than Claude Sonnet 5 ( max ), and also slightly higher than Opus 5.5 and GPT-6.

The performance also depends on the setting of “effort (level of effort)”. This adjustment will make the model spend more time thinking in exchange for better answers and higher costs. Anthropic indicates that under the settings of High effort, Sonnet can perform on par with GPT-6 Sol at FrontierCode, while the cost per individual task is about one-fifth of the latter.

In the GDPval-AA test – which uses a Elo system similar to chess rating points to score real professional jobs in 44 different professions – Sonnet scored 1844 with a 5.5, and Opus also scored 1846 with a 5.5, which can basically be considered a tie. GPT-6 and Sol scored 1487.

Competitors have also matched their prices. Last week, OpenAI reduced the price of GPT-6 and Sol to $2 per million inputs and $10 per million outputs; the mid-range model GPT-5.6 and Terra was priced at $2 per million inputs and $12 per million outputs. Anthropic has not released the benchmark test results for Terra.

The problem lies in

Sonnet 5.5 is a “high-yielding” model. Under the settings of max and effort, it generates an average of about 193,000 token per test task, which is the highest level recorded by Artificial Analysis. This is approximately 60% higher than that of Opus 5.5. Calculated in this way, the cost per task is 7.60 US dollars, which is about 50% higher than that of Sonnet 5. This does not align with the claim made by Anthropic that “up to 30% in costs can be saved”.

The savings mentioned by Anthropic come from using lower settings: under the default settings in Medium effort, the company claims that Sonnet can achieve better encoding results than the latter at less than one-tenth of the cost of the latter's best encoding performance. Artificial Analysis indicates that High effort are the most cost-effective settings. For everyday users, this means that by keeping the adjustment settings at a lower level, they can obtain nearly flagship-level encoding capabilities at a fraction of the price of a flagship model.

The table published by Anthropic consists of self-reported company data, while Artificial Analysis tested a pre-release version that contained a vulnerability. Anthropic expects that this issue will not have a significant impact on the results, or it may have merely underestimated the scores slightly. Anthropic also indicates that in complex tasks that require sustained judgment, Opus with a score of 5.5 is still significantly stronger.

Claude Haiku 5.5, designed for high-throughput, cost-sensitive applications, is expected to be launched in the coming weeks.

Tip
$0
Like
0
Save
0
Views 17
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Tesla's new Roadster release postponed due to weather conditions
Tesla announced that the outdoor launch event for the new Roadster model, originally scheduled for Thursday, has been postponed due to bad weather conditions. The new date is set for October 15th. This vehicle represents Tesla's second generation of Roadster; Musk showcased a prototype in 2017 and claimed it would go into production in 2020, but the company has missed that target several times since then.
Businessinsider
·2026-09-29 05:33:23
1
A Clever RSA Attack Outwitted a Hardware Safe – What Does This Mean for the Crypto Industry?
Researchers from the University of California, San Diego, and a French INRIA institution forged the signatures of 1,024-bit RSA keys within a hardware security module without extracting the keys themselves. They used approximately 4 billion signature requests and around 1,380 CPU core years of processing power. The study was specifically focused on RSA and did not involve the elliptic curve signatures used by Bitcoin or Ethereum. The authors stated that this does not pose an immediate threat to most modern, padded RSA implementations.
Decrypt
·2026-09-29 05:33:19
1
Michael Burry believes that the bubble of AI "may burst" sooner than he previously anticipated
In the latest investment newsletter, Michael Burry stated that he is moving up the timeline for his bearish view on the artificial intelligence boom and converting some of his short positions into put options in order to achieve more cost-effective leverage. He believes that the bubble in AI “may burst earlier than expected” and disclosed adjustments to his positions in Micron, Nebius, SOXX, and Palantir.
CNBC
·2026-09-29 05:24:42
6
US Tax Service threatens to restrict ETF tax avoidance tactics; Wall Street tax strategies under scrutiny
The U.S. Treasury Department and the Internal Revenue Service are intensifying their scrutiny of tax optimization transactions on Wall Street, involving strategies such as Box Spread ETF, swaps, and "351 conversions".
The Block
·2026-09-29 05:24:41
4
U.S. stocks closed lower
The Dow Jones fell 0.67%, the S&P 500 index fell 0.77%, and the Nasdaq index fell 0.92%. The provided content only includes scattered market updates and other miscellaneous items, lacking a complete news article, which requires manual review.
The Block
·2026-09-29 05:24:39
4
View More