Benchmarks You Can Trust
Why most ML benchmarks mislead, and how to build evaluation suites that survive contact with reality.
Seattle
Jul 17, 2024 · 4:30 PM

Speaker
Senior Research Scientist · Beacon Institute
Katherine has spent twenty years on rigorous evaluation of data and ML systems at the Beacon Institute, building the benchmarks the field actually trusts.
Why most ML benchmarks mislead, and how to build evaluation suites that survive contact with reality.
Seattle
Jul 17, 2024 · 4:30 PM
Reproducible, shareable evaluation pipelines for teams that no longer sit in the same room.
Virtual
Jul 20, 2021 · 2:30 PM