註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

VulcanBench
@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
加入 March 2020
63 正在關注    2.3K 粉絲
Announcing the next chapter.
I am excited to announce a new chapter for VulcanBench 🖖 🛡️ We are at a time in history where AI safety could possibly be most important way to make an impact in the world. As one human, I'm trying to increase the impact I can make. And I think there's an opportunity here, that I just can't stop thinking about. It's a different approach from what companies like METR are taking. And I'm not diminishing what they are doing in any way, but I am saying there is room for other approaches. I want to look at AI safety, as companies use AI today, in the actual harnesses they use, with the actual things, their teams are using AI for every single day. This is a big mission, and an important one, and in many ways, I am finding a bigger life purpose through it. More on VulcanBench Safety v1 here:
顯示更多