模型

分析 3 万场 Arena 对战后,Claude Fable 5 的概念相似度排名出人意料

We analyzed how similar model responses were across 30,086 Arena battles. Models shared 43% of thei...

精选理由

朋友分析 3 万场对战数据,发现 Claude Fable 5 的概念相似度排名和预期不同,挺有意思的。

我们分析了 30,086 场 Arena 对战,发现模型平均共享 43% 的想法。通常认为同一实验室或国家的模型概念重叠度更高,但结果并不一致。Claude Fable 5 的概念匹配度最高的是 Opus 和 Sonnet 之外的其他模型,这一发现可能令人意外。

原文 · lmarena.ai

We analyzed how similar model responses were across 30,086 Arena battles. Models shared 43% of thei...

We analyzed how similar model responses were across 30,086 Arena battles. Models shared 43% of their ideas on average. We might expect that models from the same lab, or country, would show greater conceptual overlap. But the results don’t consistently support that. Claude Fable 5 illustrates this pattern: its closest conceptual match was neither Opus nor Sonnet. Which model came closest, along with the broader findings, may surprise you. Check out the full article from @DawidGalarowicz and @petergostev below. Arena.ai @arena x.com/i/article/2098… 🔗 View Quoted Tweet 💬 10 🔄 15 ❤️ 119 👀 9032 📊 20 ⚡