註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Chris
@ChrisGPT
Agi 2029 - AI Insider / Reporter as featured in Axios • The Information • NYT • Techcrunch
加入 January 2023
1.6K 正在關注    69.4K 粉絲
There’s basically unanimous praise for Opus 5.5, but I think people *might* be overlooking what its existence implies about Anthropic’s internal models. Anthropic almost certainly has substantially stronger internal models helping generate training environments. Think of the stronger internal model as the teacher and Opus 5.5 as the cheaper deployable student. Now Opus 5.5 itself scores 55.8% on CoBench 2.1, while Anthropic estimates roughly 85% would be required to fully substitute for its research staff. Opus 5.5 is now only - 30 percentage points away from Anthropic’s benchmark threshold for fully substituting its research staff.
顯示更多
Anthropic is sandbagging btw. Just like OpenAI both have models significantly more powerful than Opus 5.5 or Astra
0
15
285
16
轉發到社區