注册并分享邀请链接,可获得视频播放与邀请奖励。

Gerrit De Vynck 🦭
@GerritD
I write about AI for the @WashingtonPost. Signal me at GerritD.27 email: gerrit.devynck@washpost.com.
加入 January 2009
7.1K 正在关注    11.1K 粉丝
more disclosures from OAI
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections
显示更多