가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Rico Angell
@rico_angell
AI Safety Researcher. Postdoc at NYU CDS.
가입 September 2024
314 팔로잉 중    146
Think your model is safe? Just ask it again and again! Language models are probabilistic. We find that just repeatedly asking the same query, with no jailbreaks, can yield harmful outputs. If your LLM is queried millions of times a day, these rare misbehaviors are inevitable!
더 보기