Wed, Oct 7, 2026 · 2:00 PM – 2:30 PM
Modern LLM applications demand infrastructure that handles massive scale reliably. We'll explore distributed serving patterns, batching strategies, and optimization techniques that work in production. Through real examples from systems processing billions of tokens daily, you'll learn how to architect inference platforms that maintain low latency while maximizing throughput. We'll dive into containerization, load balancing, and graceful degradation patterns. By the end, you'll understand the key tradeoffs in building production-grade LLM infrastructure and how to make architectural decisions that balance cost, latency, and reliability. Bring your current p99 and your monthly inference bill; the capacity model we walk through is the one that took ours from 940ms to 210ms without adding a single GPU.