註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rico Angell
@rico_angell
AI Safety Researcher. Postdoc at NYU CDS.
加入 September 2024
314 正在關注    146 粉絲
Think your model is safe? Just ask it again and again! Language models are probabilistic. We find that just repeatedly asking the same query, with no jailbreaks, can yield harmful outputs. If your LLM is queried millions of times a day, these rare misbehaviors are inevitable!
顯示更多