注册并分享邀请链接,可获得视频播放与邀请奖励。

Martian
@withmartian
Understanding Intelligence. Measurement. Explanation. Application. That's how we're tackling AI interpretability: the greatest scientific problem of our age.
加入 March 2023
14 正在关注    3.8K 粉丝
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: Academic Paper:
显示更多
0
58
219
21
转发到社区