Harvey分享开源数据集,低成本建设研究实验室的经验值得学习。
Harvey在Sovereign AI活动上分享了开源数据集和建设低成本研究实验室的经验,包括法律代理、合成数据生成、模型服务矩阵等内容。
源:https://t.co/ZpUpyNXiVm
源: x.com/gradypb/status… Pat Grady @gradypb Want world class research capabilities, but don’t have the resources of a big lab? At our recent Sovereign AI event, @gabepereyra shared @harvey ’s “moneyball” approach. Here’s the playbook: 00:00 Introduction 00:37 Building a research lab on a budget 02:28 Legal Agent Bench, contracting, and the diligence dataset 03:57 Domain experts guiding synthetic data generation 05:23 Why Harvey open sourced its datasets 06:55 Working with the neo labs – and why more than one 08:20 Post-training in-house: building "Associate 1" 09:44 The model serving matrix: 60 countries, fallbacks, SLAs 11:05 Deciding what stays in production 12:29 Simple open source switches and model routing 13:55 Moneyball: "If we win on this budget, we change the game" 14:53 Q&A: Training with sensitive data 17:16 Q&A: Competing for research talent 18:46 Q&A: Designing rubrics that actually challenge frontier models 20:19 Q&A: Where the pipeline breaks — data, research, or infra 22:59 Q&A: The tension in open sourcing a benchmark 25:02 Q&A: Biggest remaining open problems 27:10 Q&A: Competing with horizontal products Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 399 ⚡