OpenAI研究主管Mark Chen谈扩展定律与评估危机

In this episode, @OpenAI Chief Research Officer @markchen90 joins @allenpark to flambé shrimp, cook ...

精选理由

OpenAI研究老大亲口聊评估危机和扩展定律,全是干货,没有废话。

AI 摘要

OpenAI首席研究官Mark Chen在播客中讨论了扩展定律和预训练仍具重要性,解释了OpenAI如何选择研究方向和分配算力。他指出当前AI评估存在危机,并警告基准测试过拟合(benchmark-maxing)的问题。Chen还探讨了多模态推理、长期实际任务处理以及端到端AI研究的未来路径。他认为研究人员需要培养“研究品味”以避开无意义的优化。

原文 · Latent.Space

In this episode, @OpenAI Chief Research Officer @markchen90 joins @allenpark to flambé shrimp, cook ...

In this episode, @OpenAI Chief Research Officer @markchen90 joins @allenpark to flambé shrimp, cook Korean stew, and chat about being at the frontier of AI research: why scaling laws and pre-training still matter, how OpenAI chooses research bets and allocates compute, what it means to develop research taste, why evals are in crisis, how to avoid benchmark-maxing, and what it will take for models to handle long-horizon real-world work, multimodal reasoning, and eventually end-to-end AI research. Timestamps: 0:00 Intro 0:28 The Soup Story 1:52 From Trading to AI Research 3:21 How to Develop Research Taste 5:23 RL, Evals, and Superhuman Benchmarks 8:17 Cooking Begins on the Impulse Stove 8:53 Scaling Laws, Pre-Training, and Reasoning 12:33 OpenAI’s Research Roadmap and Compute Allocation 15:48 What Makes a Great Researcher 19:33 The Evals Crisis and Benchmark-Maxing 24:34 Jagged Intelligence, Context, and Long-Horizon Learning 27:14 Shrimp Flambé and New Research Bets 31:32 Multimodal Models and One Architecture 32:36 Vibe Researching and End-to-End AI Research 34:36 Failed Bets, Postmortems, and OpenAI’s Alpha 37:07 Final Taste Test 37:53 Overrated vs. Underrated AI Research 41:00 Closing Your browser does not support the video tag. 🔗 View on Twitter 💬 4 🔄 1 ❤️ 34 👀 5289 📊 16 ⚡