callboard

Frontier AI Summit 2026

← Back to the speaker gallery

Marcus Chen

Senior Eval EngineerClarity Evals

I build evaluation frameworks for LLMs at Clarity. My background spans both ML research and production systems. I focus on designing evals that catch real-world failures without becoming brittle benchmarks. Previously worked on safety systems at several major labs. Believe in making eval tooling accessible to all developers. I gave the eval-drift talk at NeurIPS last December and have been arguing about the methodology in DMs ever since.Show more

I build evaluation frameworks for LLMs at Clarity. My background spans both ML research and production systems. I focus on designing evals that catch real-world failures without becoming brittle benchmarks. Previously worked on safety systems at several major labs. Believe in making eval tooling accessible to all developers. I gave the eval-drift talk at NeurIPS last December and have been arguing about the methodology in DMs ever since.

Sessions (1)

View full schedule