注册并分享邀请链接,可获得视频播放与邀请奖励。

Rico Angell
@rico_angell
AI Safety Researcher. Postdoc at NYU CDS.
加入 September 2024
314 正在关注    146 粉丝
Think your model is safe? Just ask it again and again! Language models are probabilistic. We find that just repeatedly asking the same query, with no jailbreaks, can yield harmful outputs. If your LLM is queried millions of times a day, these rare misbehaviors are inevitable!
显示更多