登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Alex Shaw
@alexgshaw
Hacking on @terminalbench and @harborframework. Founding MTS @LaudeInstitute. Formerly Google. BYU alum.
参加 October 2021
771 フォロー中    2.8K ファン
I’ve been thinking more about this and wanted to share my thoughts. I’m glad Epoch is focusing on eval quality, but I think their approach is flawed. The flaw is made clear by the fact that 4 benchmarks weren’t labeled “flawed”. The fact of the matter is, benchmarks are software and all software has bugs. Would you label nextjs as categorically flawed if you found a bug in it? This is why we’re pushing the industry towards continuous benchmarks. Our approach to addressing benchmark bugs is not to tweet a binary “flawed vs verified” label, but instead provide the tools for anyone to create, improve, and maintain benchmarks, while easily and cost-effectively reconciling their results to the latest version. We would love to work with the Epoch team to encode some of their verification practices into tools that people can use to continuously improve their benchmarks.
もっと見る
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
もっと見る