We have something very special to share today!
We developed a standard to assess benchmark quality, which we plan to apply to the most-used AI capability metrics going forward.
We hope this will raise the bar for designing and interpreting evaluations. Let us know what you think!