Wed, Oct 7, 2026 · 6:00 PM – 6:30 PM
Standard benchmarks often miss the failures that matter in production. This talk explores building comprehensive evaluation frameworks that actually predict real-world performance. We'll discuss what makes evals brittle versus robust. You'll learn techniques for designing evals that are rigorous without becoming expensive or gaming-prone. We'll cover automated evaluation systems, human-in-the-loop approaches, and continuous monitoring. Practical examples from production systems show how different evaluation strategies catch different failure classes. Walk away with actionable patterns for building better evals. Fair warning: about half of this talk is failures — three eval suites we shipped that stayed green while the model got measurably worse, and what we changed to make them able to fail.