Arena 推出评估产品8个月后年收入达1亿美元

Arena reached a $100M annual revenue run rate just 8 months after launching our evaluation product. ...

精选理由

Arena 8个月做到1亿美元年收入,它的 Agent Arena 能测 AI 智能体在真实任务里的表现,比传统投票评测更硬核。

AI 摘要

Arena 是一个从 UC Berkeley 研究项目起步的 AI 评估平台。推出评估产品仅8个月后,其年化收入运行率突破1亿美元。平台推出 Agent Arena,用于评估长期运行的智能体在复杂现实任务中的表现,包括工具使用、任务完成率和幻觉率。目前 Arena 拥有数千万用户。

原文 · lmarena.ai

Arena reached a $100M annual revenue run rate just 8 months after launching our evaluation product. ...

Arena reached a $100M annual revenue run rate just 8 months after launching our evaluation product. We started as a research project at UC Berkeley with a simple mission: measure AI progress through real-world use. As AI shifts from chatbots to agents taking on longer, higher-stakes work, the problem matters more than ever. Today, Arena measures real-world AI utility with a community of tens of millions. With Agent Arena, we’re evaluating long-running agents on complex, real-world tasks - how they use tools, adapt to feedback, recover from errors, and accomplish goals set by humans. We are excited to keep deepening our work in agentic evaluations. Here’s @ml_angelopoulos on what this milestone means and where we go from here: Your browser does not support the video tag. 🔗 View on Twitter Anastasios Nikolas Angelopoulos @ml_angelopoulos Arena has crossed $100M in annualized revenue run rate, eight months after launching our evaluation product. With our recent release of Agent Mode, millions of users on Arena are doing real work with agents, from coding to document analysis, in long-running, multi-turn sessions with hundreds of tool calls. Arena now evaluates objective criteria like task completion rates, hallucination rates, and more, far beyond our original human preference voting model. This expansion has taken us from a student project at Berkeley to one of the fastest growing companies in history. Go Bears! 🐻 Our core thesis is simple: to align AI with human values, we must directly measure its impact on people in the real world. Today's milestone is proof that Arena’s platform is the de-facto standard for post-deployment evaluation of AI. 🔗 View Quoted Tweet 💬 32 🔄 31 ❤️ 370 👀 145576 📊 80 ⚡