Arena 的数据管道设计解决了大规模 AI 评估元数据管理的痛点,做评测平台或数据管线的团队可以直接借鉴其可插拔标签框架和成本控制思路。
Arena 研究人员 Guanglei Song 和 I-Hung Hsu 在视频中详细介绍了 Arena 分类排行榜背后的数据管道:从 Databricks 和 Spark 作业到可插拔标签框架,调用 LLM 对文本、图像、前端编码等领域的每次评估进行分类。这个元数据层让 Arena 数据超越排行榜排名,对研究更有用。视频还涵盖了动态并发控制处理不稳定的 LLM API、无需重建系统即可添加新标签器、以及成本控制策略(过滤、幂等性和模型选择)。
Millions of votes a week. One tagging system. Arena researchers Guanglei Song and I-Hung Hsu walk t...
Millions of votes a week. One tagging system. Arena researchers Guanglei Song and I-Hung Hsu walk through the data pipeline behind Arena's category leaderboards: Databricks → Spark → a pluggable tagger framework calling LLMs to categorize every evaluation across our text, image, frontend coding, and other arenas. This metadata layer is what makes Arena data useful for research beyond just leaderboard rankings. 0:00 How Arena collects evaluation data 1:50 Pipeline architecture: Databricks and hourly Spark jobs 2:35 The pluggable tagger framework 4:35 Handling flaky LLM APIs with dynamic concurrency control 6:30 Adding new taggers without rebuilding the system 7:30 Backfilling history alongside the live stream 9:10 Cost control: filtering, idempotency, and model selection 11:10 Chunking long messages Your browser does not support the video tag. 🔗 View on Twitter 💬 2 🔄 6 ❤️ 55 👀 6940 📊 9 ⚡