註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Martian
@withmartian
Understanding Intelligence. Measurement. Explanation. Application. That's how we're tackling AI interpretability: the greatest scientific problem of our age.
加入 March 2023
14 正在關注    3.8K 粉絲
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: Academic Paper:
顯示更多
0
58
219
21
轉發到社區