Artificial Analysis 两周年:从 o1-preview 到覆盖智能体任务的评价体系
Artificial Analysis 发文回顾两年:o1-preview 开创推理模型,他们的评测指数也从 4 项考试题扩到 10 项智能体任务,能看清评测怎么进化。
Artificial Analysis 回顾两周年,OpenAI 在两年前发布 o1-preview,是首个通过推理 token 先思考再作答的推理模型,如今所有前沿模型都采用这一方式。Artificial Analysis Intelligence Index v1 当时只用 MMLU、GPQA、MATH、HumanEval 四项单轮考试式评测。目前最新的 v4.3 版本已纳入 10 项更难的评测,覆盖长程智能体任务、高难度编程和知识型工作。
Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to ‘think’ before answering
Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.