行业79°

Anthropic产品负责人Dianne Penn的六大核心心得

My biggest takeaways from @AnthropicAI's Head of Product for AI research and Labs, Dianne Penn: 1. ...

精选理由

Anthropic产品老大亲述Claude Code为何离不开Opus 4.5,以及eval怎样取代PRD成为核心工作流,挺有启发的。

AI 摘要

Anthropic产品负责人指出,Opus 4.5和Claude Code相互成就,模型智能与产品载体交汇带来爆发时刻。团队将eval视为产品工作的核心工件,每个新功能生成30-40个代表性示例编码为带标准答案的提示词。产品工作需像关注像素一样关注token,通过分析AI transcript诊断失败是幻觉、过度自信还是工具调用错误。2024年初的Golden Gate Claude实验在24小时内上线,虽仅触达2000人,但证明了Anthropic能提供差异化体验。AI能力呈离散跳跃,没有平滑过渡,需要eval系统来捕捉突然的能力跃升。

图片来源 · Lenny Rachitsky
原文 · Lenny Rachitsky

My biggest takeaways from @AnthropicAI's Head of Product for AI research and Labs, Dianne Penn: 1. ...

My biggest takeaways from @AnthropicAI 's Head of Product for AI research and Labs, Dianne Penn: 1. You need frontier products to feel frontier models. Opus 4.5 wouldn’t have had its breakout moment without a vehicle like Claude Code—and Claude Code wouldn’t have seen that acceleration in adoption without Opus 4.5. The team had felt Claude Code’s magic internally for months; the inflection came when model intelligence and product vehicle finally met. 2. Evals are the new PRDs. Anthropic’s research PM team treats the eval set as the primary artifact of product work. A PRD describes what to build; an eval defines what success looks like once you’ve built it. For every major new feature, Dianne and her team generate 30 to 40 representative examples and encode them as prompts with golden answers. PRDs still exist, though, particularly for aligning large groups across engineering, legal, and safety, but the core of her product work has become the eval. 3. Sweat the tokens as much as you sweat the pixels. Product work has long been obsessing over every pixel in a user flow. More and more, the craft is reading AI transcripts, studying failed trajectories, and diagnosing whether a miss was a hallucination, an overconfident answer, or a wrong tool call. Each diagnosis routes to a different team and a different fix, and getting to that level of detail is becoming the core of the job. 4. "Golden Gate Claude" was an inflection point for Anthropic. In early 2024, Anthropic’s research team identified “features” inside the model’s layers—latent themes like bullet-point writing or geographic landmarks. Researchers dialed up the 'Golden Gate Bridge' feature, and Claude became obsessed with it, weaving the bridge into every response, including a spaghetti recipe. The team spun up a consumer experience at claude.ai within 24 hours—engineering, product, and design donating their time off. It reached just 2,000 people. But it proved Anthropic could ship experiences that were different from competitors and authentic to its research. “That was one of those hidden inflection points of finding our identity.” 5. AI capabilities arrive discontinuously—and evals are how you catch them. The original scaling law papers have two sets of graphs: the smooth curve of loss going down as you add compute and data, and jagged graphs of emergent capabilities. Models go from being unable to calculate 1 + 1 to calculating it reliably every time, with no smooth transition. Nobody knows the exact moment a jump will happen—“unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don’t know.” It’s also what makes safety harder. And it means there’s real “product overhang and user overhang” sitting in today’s models, waiting to be discovered. 6. AI writing is the current jagged edge—and it’s being actively sharpened. Why is AI-written text still so recognizable? The technology is jagged: the last push made models dramatically more agentic, which made writing the new rough edge. There are now active efforts on her team and the research side to make Claude write better, including tone and character. 7. Anthropic’s Labs team exists to make discontinuous 10x–1,000x bets. Labs’ thesis is pulling on threads that aren’t on the core roadmap and asking what the 10x, 100x, 1,000x version is—work that produced Claude Code, MCP, Skills, and Claude Design. The pods are tiny; ideas sometimes start with one engineer, because “really large teams pursuing very ambiguous, large ideas end up being slowed down.” Bets that don’t work yet get revisited one to two model generations later, and prototypes that never ship still count—they become learning about where the models are strong. The hard part is people: you select for founder types who can pour their heart into a bet, watch it get turned off, and go again. 8. We will need PMs more than ever. Dianne asked this of herself—do we still need PMs when models are this capable and engineers are able to build anything? Her answer: “The role of people who are user-centric, who go into the details of understanding what users are trying to accomplish, bubbling that up in an actionable manner, and doing the relentless work to do that—that to me is the core of a product person, and we actually need more of that.” As AI makes it easier to build, the scarce human skill becomes judgment—which she describes as an accumulation of nuance and experience the models haven’t lived yet. 9. Managers who aren’t hands-on building are losing their theory of mind. Dianne’s onboarding plan for tenured PMs is identical to the one for early-career hires—reading transcripts, talking to users, understanding what good looks like. The reason: you cannot make good product decisions about AI features if you’re not talking to the model regularly. She carves out one to two active workstreams on every model release to keep her own mental model of how capabilities are moving. “If you’re not building yourself, you’re not gonna make it.” 10. Protect your brain from AI brain rot. To avoid over-relying on AI, form your point of view first, then use Claude as a sparring partner. Delegate fully only where the writing matters less than the thinking (like a standardized monthly business review) and shift into the reviewer’s seat—because who signs off matters more than who wrote it. 11. You can use AI to raise your EQ, not just your IQ. Dianne built a custom Claude “skill” based on the book Crucial Conversations that coaches her before high-stakes moments—helping her go deeper faster, build trust, and be more direct—and she shares it with other managers on her team. 12. Alignment makes Claude more useful, not less. People assume safety work limits capability; Dianne’s experience is the opposite. Anthropic’s theory of change is that reducing sycophancy—the tendency to agree with whatever you present—produces improvements in judgment and output. She’s used a research version of Opus to pressure-test Claude pricing decisions (yes, using Claude to price Claude), specifically because it pushes back. “A thinking partner doesn’t just agree with you. It should add to you, and you should come away having better ideas because you worked with Claude. 13. Claude’s coding dominance began as a relatively small training change. In 2023, the company was under 200 people, and up to that point, “nobody said Anthropic and coding in the same sentence.” GPT-4 was used a bit for coding, but mostly for autocomplete-style code. The team noticed people writing long-form code and decided to train Opus 3 to be better at it. It was a relatively small change to the training run but differentiated Claude in the market and won its earliest developer enthusiasts. The bet was also strategic: Anthropic’s core thesis of recursive self-improvement requires agentic models that write code and use tools well. Lenny Rachitsky @lennysan Dianne Penn is Head of Product for Research and Labs at @AnthropicAI . She joined in 2023 as their first technical PM (when there were just five product engineers) and has helped ship every Claude model from Claude 2 to Mythos. We discuss: 02:31 Early Anthropic 08:55 The two biggest inflections 13:50 Inside the exponential 20:02 Token maxing 23:30 Anthropic Labs and the incubation model 27:30 How the AI research role works 31:35 How to become a top AI researcher 35:18 Frontier model safeguards 39:38 Hiring in the AI era 44:16 Building evals 47:48 Evals are the new PRDs 49:55 The importance of hands-on leadership 52:46 Finding joy in AI 58:10 How Dianne uses Claude 01:01:05 Avoiding overreliance on AI 01:03:50 The constitution that makes Claude better 01:07:11 AI writing and verification 01:11:40 Where human brains will continue to be valuable 01:14:10 Navigating AI with kids 01:16:26 Alignment, the future of the PM role, and burnout 01:21:54 Lightning round and final thoughts Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 6 🔄 2 ❤️ 23 👀 4560 📊 8 ⚡

Anthropic产品负责人Dianne Penn的六大核心心得 · AI 热点