行业精选

LangChain创始人谈智能体三要素:框架、模型与上下文

talked about how harnesses & evals help you own your intelligence how they fit in to the big pictur...

精选理由

红杉活动上,LangChain创始人聊了智能体三件套:框架、模型、上下文。他说越偏离模型训练数据,越该自己造框架。还演示了LangSmith Engine怎么做评估。

AI 摘要

LangChain创始人Harrison Chase在红杉活动上提出智能体由框架、模型和上下文三部分组成。他认为若使用场景偏离模型训练分布,越应自建框架。他验证了LangSmith Engine的评估功能,并展示了如何通过追踪、数据飞轮优化性能。视频中还讨论了为何可观测性常被低估。

原文 · Harrison Chase

talked about how harnesses & evals help you own your intelligence how they fit in to the big pictur...

talked about how harnesses & evals help you own your intelligence how they fit in to the big picture: owning your intelligence means three things: - open agent system (harness is a big part of this!) - compounding loop (evals are a big part of this!) - governed runtime (harness also important here - we see managed harnesses growing rapidly) Sonya Huang 🐥 @sonyatweetybird An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. @LangChain founder @hwchase17 joined us at our @sequoia Own Your Intelligence to talk about the piece that often gets the least attention: the harness. He offers a clear heuristic for when to build your own. The more out of distribution you are from what the models were trained on, the more you'll want to customize. And good technical content on how to actually measure performance with evals and langsmith. 00:00 Introduction 00:58 The three parts of an agent: harness, model, context 02:12 What a harness actually does 03:25 Customizing the core loop with middleware 04:41 Sandboxes, file systems, sub-agents, summarization 05:47 Cognitive architectures — and when you still need them 07:03 Build your own harness or use off the shelf? 08:24 In-distribution vs. out-of-distribution: the file-editing example 09:39 Why evals define what "good" means in an organization 11:04 Harbor: what an eval task actually looks like 12:11 Comparing harnesses and models on accuracy, latency, and cost 13:20 Why observability is underrated — it's usually the context 14:34 The data flywheel: traces → curation → experiments 15:42 Getting feedback through UX design and online evaluators 16:51 Demo: LangSmith Engine 19:23 Q&A: Running Engine on Engine, and "codex-ification" 20:44 Q&A: Will harnesses converge or diverge? Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 4 ❤️ 9 👀 2441 📊 3 ⚡