听听OpenAI首席研究官Mark Chen聊预训练为啥没过时、评估危机怎么破,还有未来的研究路线图,很实在的讨论。
OpenAI首席研究官Mark Chen在播客中明确表示预训练并未过时,扩展律仍然有效。他讨论了基准测试过度优化导致的评估危机,以及OpenAI如何通过新的工程和研究洞察突破边界。他还提到模型需要处理长期现实世界任务、多模态推理,最终实现端到端AI研究。
Is pre-training dead? @OpenAI Chief Research Officer @markchen90 doesn't think so: "We've always fo...
Is pre-training dead? @OpenAI Chief Research Officer @markchen90 doesn't think so: "We've always found some kind of technique whether it be better engineering or some new research insight that helps you break past the boundary." Your browser does not support the video tag. 🔗 View on Twitter Latent.Space @latentspacepod In this episode, @OpenAI Chief Research Officer @markchen90 joins @allenpark to flambé shrimp, cook Korean stew, and chat about being at the frontier of AI research: why scaling laws and pre-training still matter, how OpenAI chooses research bets and allocates compute, what it means to develop research taste, why evals are in crisis, how to avoid benchmark-maxing, and what it will take for models to handle long-horizon real-world work, multimodal reasoning, and eventually end-to-end AI research. Timestamps: 0:00 Intro 0:28 The Soup Story 1:52 From Trading to AI Research 3:21 How to Develop Research Taste 5:23 RL, Evals, and Superhuman Benchmarks 8:17 Cooking Begins on the Impulse Stove 8:53 Scaling Laws, Pre-Training, and Reasoning 12:33 OpenAI’s Research Roadmap and Compute Allocation 15:48 What Makes a Great Researcher 19:33 The Evals Crisis and Benchmark-Maxing 24:34 Jagged Intelligence, Context, and Long-Horizon Learning 27:14 Shrimp Flambé and New Research Bets 31:32 Multimodal Models and One Architecture 32:36 Vibe Researching and End-to-End AI Research 34:36 Failed Bets, Postmortems, and OpenAI’s Alpha 37:07 Final Taste Test 37:53 Overrated vs. Underrated AI Research 41:00 Closing Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 1 🔄 1 ❤️ 22 👀 4989 📊 5 ⚡