註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

AI Dance
@AI_Whisper_X
China AI insider | Silicon Valley Decoded 一边盯硅谷,一边扒中国AI 算法 + VC 双视角 · 讲人话 📬 aidance.info@gmail.com
加入 October 2024
293 正在關注    5.9K 粉絲
绷不住了,Anthropic也学会了Benchmark作弊?曾经grok玩的”数据科学“那一套,现在Anthropic也开始玩了? FrontierCode 这一行,Opus5的 53.4 > 53.5 是什么鬼?(53.5是Fable5) 然后Agentic Coding,GPT5.6最高分,给了个浅色?真是笑死我了。。。 这张图如果是人类做的,感觉 Anthropic 已经没落至此了吗?卧槽,那我当年真是爱错人了…… 这张图如果是 AI 做的,那我只能说,AI 的幻觉和谎言还是在的。 Anthropic还是太权威了。
顯示更多
look at how they highlighted the one benchmark that gpt 5.6 sol won. so dishonest