产品精选

a16z 投资 Preference Model,并开源 RL 环境框架 Karotte

We're excited to invest in Preference Model. Building RL environments that actually work is harder ...

精选理由

a16z 投了 Preference Model,这家人专门给实验室做 ML 工程的 RL 环境,还把用了一年、跑过百万次评估的框架 Karotte 开源了,做 RL 训练的可以看看。

a16z 宣布投资 RL 环境公司 Preference Model。该团队过去一年为前沿实验室构建 ML 工程方向的 RL 环境,专门对抗模型的 reward hacking 和漏洞利用。本周他们将生产环境使用的框架 Karotte 开源,该框架经过超过 100 万次评估运行和红队测试打磨。Karotte 能定位模型薄弱点、生成针对性任务,并测试环境抵抗智能体攻击的能力。

原文 · a16z

We're excited to invest in Preference Model. Building RL environments that actually work is harder ...

We're excited to invest in Preference Model. Building RL environments that actually work is harder than it looks. Models are relentless at reward hacking, finding shortcuts, and exploiting vulnerabilities. @preferencemodel has focused on the domain that matters most to labs right now: AI research and ML engineering itself. Over the past year, the team has built RL environments for leading labs. Their focus has been on building the infrastructure to make harder, more resistant environments as models improve: tooling that finds where models are weak, generates new tasks to target those gaps, and tests environments against agents actively trying to break it. This week they're open-sourcing Karotte, the framework they've used in production — hardened through more than a million evaluation runs and red-teaming. We're thrilled to partner with @chem_safety and @Ning_Catsnail and the Preference Model team as they build the training grounds for capable and aligned models. By @JenniferHli Jennifer Zhou @chem_safety Today we are open-sourcing 🥕Karotte, our framework for building RL environments. We've used it for the past year to build MLE RL environments for frontier labs, and it's been hardened through 1M+ evaluation runs and red-teaming. 🔗 View Quoted Tweet 💬 5 🔄 9 ❤️ 65 👀 8201 📊 10 ⚡