注册并分享邀请链接,可获得视频播放与邀请奖励。

Xiangyu Qi
@xiangyuqi_pton
Cyber Models @openai | PhD @Princeton | Prev @GoogleAI @GoogleDeepMind
加入 December 2019
1.1K 正在关注    4K 粉丝
Astra achieves a full 100% success rate on ExploitBench, so we had to build an internal refresh using newly disclosed vulnerabilities from June through August that fall after the model’s knowledge cutoff. On this refreshed benchmark, Astra remains dramatically stronger than GPT-5.6 Sol while using far fewer tokens. This result, together with several other pieces of evidence, has led us to believe that Astra has reached the “cyber-critical” capability threshold under our Preparedness Framework.
显示更多
0
62
1.9K
140
转发到社区