模型多源确认88°

TasteVal 基准发布:量化 AI 的实验研究品味,Opus 5.5 超过人类专家

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

精选理由

这个 TasteVal 把研究品味换算成省了多少算力来打分,Opus 5.5 用 1/30 的成本跑赢人类专家,数字挺有意思的。

新基准 TasteVal 用 8 个开放式任务评测前沿模型的实验研究品味,即给定固定研究问题后迭代设计实验并解读结果的能力。被测模型扮演 Researcher,由固定的 Coder agent 负责实现实验,预算上限为 40 H100 小时或 120 小时。团队招募 24 名人类专家建立基线,评测了 2023 至 2026 年发布的 20 个模型。成绩最好的 Opus 5.5 超过专家基线,计算乘数为 2.3 倍(95% CI 1.15-4.37),单次运行成本约为人类基线的 1/30。数据还显示自 2025 年 12 月起,前沿模型的计算乘数约每 3.0 个月翻一倍,明显快于此前的 14 个月周期。

原文 · arXiv cs.AI

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15-4.37), at roughly 1/30 of our baseliners' average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7-5.0), up from every 14 months between 2023 and December 2025. Measured by final normalized performance, frontier models show no trend break, doubling every 14.6 months. To keep TasteVal uncontaminated, we do not release the tasks.