Anthropic 发布 Haiku 5.5,成本降低 90%,OSWorld 得分升至 72.4%
Anthropic 的小模型 Haiku 5.5 便宜了 90%,操作电脑的基准分从 15.7% 拉到 72.4%,跑批量智能体任务的可以看看。
Anthropic 发布 Haiku 5.5,在 10 万 token 以内的提示词场景下,成本比 Haiku 4.5 低 90%。该模型首次引入 Haiku effort 设置,可在成本与准确率之间权衡。在测试智能体操作真实电脑能力的 OSWorld 2.1 基准上,得分从 Haiku 4.5 的 15.7% 提升到 72.4%;在无工具图表推理测试 Chartography 上从 6.4% 升至 46.4%。Asana 实测任务完成延迟降低超 30%,每轮智能体推理速度提升至 2.5 倍。不过 Anthropic 仍建议 Terminal-Bench 4.0 等复杂智能体编程任务使用 Sonnet 5.5 或 Opus 5.5。
Anthropic dropped Haiku 5.5
> costs 90% less than Haiku 4.5 on prompts up to 100K tokens.
> Haiku 5.5 adds the first Haiku effort setting, trading cost for accuracy, but Anthropic still recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding such as Terminal-Bench 4.0 tasks.
> Asana reported over 30% lower task-completion latency and up to 2.5x faster inference per agent turn than its current model.
> On OSWorld 2.1, which tests whether an AI agent can operate a real computer to finish long multi-step tasks, Haiku 5.5 jumped from Haiku 4.5's 15.7% to 72.4%.
On Chartography, a visual reasoning test of reading and interpreting charts without tools, it rose from 6.4% to 46.4%.