Muse Spark 1.3重回WebDev编程榜前十

Muse Spark 1.3 (xHigh) just landed @AIatMeta back in the top 10 models for Code Arena: WebDev! This...

精选理由

Meta的Muse Spark 1.3在编程榜单上跃升20多位,性能接近Claude和GPT-5.6,工具调用减少20%。

AI 摘要

Muse Spark 1.3 (xHigh)在Code Arena: WebDev排行榜上位列第10,得分为1623分(AutoEval)。该模型性能与排名第8的Claude Fable 5(1628分)和排名第11的GPT-5.6 Sol(1616分)相当。相比前代模型,Muse Spark 1.1排名第25(1540分),Muse Spark 1.2 (xHigh)排名第30(1534分),新版本有明显提升。

原文 · lmarena.ai

Muse Spark 1.3 (xHigh) just landed @AIatMeta back in the top 10 models for Code Arena: WebDev! This...

Muse Spark 1.3 (xHigh) just landed @AIatMeta back in the top 10 models for Code Arena: WebDev! This release is ~ #10 in Code Arena: WebDev with 1623 pts (AutoEval). It’s performance is par with Claude Fable 5 at #8 (1628 pts) and GPT-5.6 Sol at #11 (1616 pts). Muse Spark 1.3 (xHigh) is a huge improvement from previous models Muse Spark 1.1 at #25 (1540 pts), and Muse Spark 1.2 (xHigh) at #30 (1534 pts). Note: this is an early AutoEval score, in which a Reward Model trained on Arena’s human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. Congrats to @alexandr_wang and the @AIatMeta team on this release! AI at Meta @AIatMeta We’re excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability. Key capabilities: → Sustains longer-horizon work across multiple workflows in a single thread → More actively collaborates with users: it asks clarifying questions, flags when it's stuck, confirms before consequential actions → Better calibrated on its own limits instead of hallucinating outcomes → ~20% fewer tool calls and ~25% fewer tokens vs. Muse Spark 1.2 in internal comparisons 🔗 View Quoted Tweet 💬 9 🔄 1 ❤️ 129 👀 13581 📊 17 ⚡