Back to issue
watch1h 34m

How Anthropic builds products like Claude Code before the AI models are ready | Dianne Penn

Dianne Penn (Head of Product, Research & Labs, Anthropic) · Lenny's Podcast

A rare inside look at how Anthropic builds products before the models can support them, from the person running its labs teams. The most useful idea for builders is that evals have replaced PRDs: PMs now express user needs as evaluation frameworks researchers can act on directly. Penn also explains the co-dependency between Claude Code and Opus 4.5 — neither would have landed without the other.

  • Evals are the new product spec: instead of writing requirements documents, PMs encode user needs into evaluation frameworks that model researchers can train against.
  • 'Sweat the tokens as much as the pixels' — reading raw conversation transcripts is how you diagnose whether a failure is tool use, retrieval synthesis, or an alignment gap.
  • Frontier products and frontier models need each other: Opus 4.5 succeeded partly because Claude Code existed as the vehicle to showcase it.
  • Anthropic plans for forward compatibility by asking 'what changes when Claude 8 arrives?' so today's product decisions don't become obsolete.
Watch on YouTube

Part of Issue Nº 003: Evals Are the New PRDs: How Anthropic, Bridgewater, and SonderMind Actually Ship AI