← Back to live feed

Monday, Sep 28, 2026

1
Claude Opus 5.5 Debuts at No. 2 in Agent Arena at 64% Lower Cost Per TaskPVT:ANTH
topics 🤖 AI tags AIAI ModelsAI ReleasesAI Research PVT:ANTH keywords Anthropic

Claude Opus 5.5 debuted at No. 2 in Agent Arena with a net improvement score of +12.15%, behind only Fable 5.1 (Max), while its $1.31 median price per task came in at 64% less cost and pushed out the Pareto frontier. The High setting posts a higher net improvement score than both prior Opus 5 variants while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, it ranks first in steerability (+14.50%), second in confirmed success (+15.50%), third in praise versus complaint (+19.80%) and fourth in bash recovery (+10.64%).

The result extends a run of leaderboard entries for the model, which also ranks No. 1 on Xbench with +76. A faster Sonnet 5.5 is being grey-tested in Claude Code at $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20.

Anthropic released Opus 5.5 on Sept. 22 as the first model in its Claude 5.5 family, priced at $4/$20 per million input/output tokens, down 20% from Opus 5. The company says it matches Fable 5.1 on most tasks and costs 40% less to run than Opus 5 on typical workloads, with output generated more than 30% faster. Sonnet 5.5 and Haiku 5.5 are due in the coming weeks.

Image via @claudeai on X
Continued in
Anthropic Releases Claude Sonnet 5.5, More Than 30% Faster and Up to 30% Cheaper Than Sonnet 5
192 tweets • 109 sources
Continues from Wednesday, Sep 23
Anthropic’s Claude Opus 5.5 Takes No. 1 in Code Arena WebDev
135 tweets • 89 sources
See all 136 tweets →