註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xiangyu Qi
@xiangyuqi_pton
Cyber Models @openai | PhD @Princeton | Prev @GoogleAI @GoogleDeepMind
加入 December 2019
1.1K 正在關注    4K 粉絲
Astra achieves a full 100% success rate on ExploitBench, so we had to build an internal refresh using newly disclosed vulnerabilities from June through August that fall after the model’s knowledge cutoff. On this refreshed benchmark, Astra remains dramatically stronger than GPT-5.6 Sol while using far fewer tokens. This result, together with several other pieces of evidence, has led us to believe that Astra has reached the “cyber-critical” capability threshold under our Preparedness Framework.
顯示更多
0
62
1.9K
140
轉發到社區