模型多源确认83°

Claude Sonnet 5.5 发布,Terminal-Bench 4.0 得分提升至 70.6%

精选理由

Anthropic 推出 Sonnet 5.5,价格不变但性能大增,成本降三成,速度提三成,特别适合编程和文档工作。

Anthropic 发布 Claude Sonnet 5.5,在 Terminal-Bench 4.0 基准测试中得分从 10.3% 提升至 70.6%。新版本任务成本降低 30%,输出速度提升 30% 以上。Sonnet 5.5 在低或中等努力设置下,能以约十分之一的成本超过 Sonnet 5 的最高分。

原文 · rohanpaul_ai

Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices.

Overall, 30% cost reduction per-task due to faster speeds and fewer tool calls.

Keeps Sonnet 5's $2/$10 per million input/output tokens, half Opus 5.5's rates.

Its savings instead come from doing less work per job, since Anthropic says fewer tokens and tool calls cut the total cost of a task by up to 30%.

Output also arrives more than 30% faster than Sonnet 5's, making Sonnet 5.5 Anthropic's quickest Sonnet yet.

Users can dial an effort setting, trading longer reasoning and more self-checking for a higher cost per task.

The economics might count for more than the leaderboard. Sonnet 5.5 operating at Low or Medium effort is able to exceed Sonnet 5's top score at about one-tenth the cost per task. On FrontierCode, it says Sonnet 5.5 at High effort scores roughly 10 points above Sonnet 5 at the same setting while costing approximately one-fifteenth as much per task.

Anthropic characterizes Sonnet 5.5 as ideal for comparatively well-defined routine work, such as software debugging, coding, document production, building presentations and spreadsheets, and designing or refining interfaces.