Evaluating LLMs You Can Actually Trust
A practical framework for building evaluation suites that catch real regressions, not vibes.
Seattle
Jul 17, 2024 · 9:00 AM

Speaker
Senior AI Researcher · Beacon Institute
Aisha researches responsible AI and rigorous model evaluation at the Beacon Institute, advocating for measurable safety and fairness in deployed systems.
A practical framework for building evaluation suites that catch real regressions, not vibes.
Seattle
Jul 17, 2024 · 9:00 AM