← Back to speakers
Marcus Chen
Senior Eval EngineerClarity Evals
I build evaluation frameworks for LLMs at Clarity. My background spans both ML research and production systems. I focus on designing evals that catch real-world failures without becoming brittle benchmarks. Previously worked on safety systems at several major labs. Believe in making eval tooling accessible to all developers. I gave the eval-drift talk at NeurIPS last December and have been arguing about the methodology in DMs ever since.Show moreShow less
I build evaluation frameworks for LLMs at Clarity. My background spans both ML research and production systems. I focus on designing evals that catch real-world failures without becoming brittle benchmarks. Previously worked on safety systems at several major labs. Believe in making eval tooling accessible to all developers. I gave the eval-drift talk at NeurIPS last December and have been arguing about the methodology in DMs ever since.