七项研究入选2026春季AI评估项目

Congratulations to seven research projects that were selected for the Spring 2026 cohort. - Emmanue...

精选理由

斯坦福、MIT等七校研究团队入选2026春季AI评估项目,聚焦LLM评估方法和智能体评估框架。

AI 摘要

斯坦福大学Emmanuel Candès研究统计高效LLM评估与排名。芝加哥大学Haifeng Xu研究LLM预测智能的系统性评估。麻省理工学院Negin Golrezaei开发多轮AI评估的因果推理框架。卡内基梅隆大学研究如何防止排行榜上的Goodhart定律被利用。

图片来源 · lmarena.ai
原文 · lmarena.ai

Congratulations to seven research projects that were selected for the Spring 2026 cohort. - Emmanue...

Congratulations to seven research projects that were selected for the Spring 2026 cohort. - Emmanuel Candès, Stanford University: Statistically Efficient LLM Evaluation and Ranking with Pairwise Human Preference Data - Haifeng Xu @haifengxu0 , University of Chicago: Towards a Principled Evaluation of LLMs’ Predictive Intelligence - Lihua Lei @lihua_lei_stat , Stanford University: Is Bradley–Terry Enough? Heterogeneous Preference Modeling for RLHF - Negin Golrezaei @NeginGolrezaei , Massachusetts Institute of Technology: Causal inference framework for multi-turn AI evaluation. - Ramesh Raskar @raskarmit , Massachusetts Institute of Technology: Toward Adaptive, Task-Conditional Evaluation for AI Agents - Steven Wu @zstevenwu and Andrew Ilyas @andrew_ilyas , Carnegie Mellon University: Goodhart’s Law on the Leaderboard: Can Cheap Preference Optimization Game Arena? -Tal Linzen @tallinzen , New York University: A Rigorous Evaluation of Meta-Cognitive Monitoring in LLMs Read more about the program details and submit your proposal for Fall 2026 here: arena.ai/blog/academic-… 💬 0 🔄 0 ❤️ 3 👀 1725 📊 1 ⚡

七项研究入选2026春季AI评估项目 · AI 热点