Register and share your invite link to earn from video plays and referrals.

Rico Angell
@rico_angell
AI Safety Researcher. Postdoc at NYU CDS.
314 Following    146 Followers
Think your model is safe? Just ask it again and again! Language models are probabilistic. We find that just repeatedly asking the same query, with no jailbreaks, can yield harmful outputs. If your LLM is queried millions of times a day, these rare misbehaviors are inevitable!
Show more
It’s deployment time! You’ve done the pre-deployment evals. You THINK your model is safe, so you ship it 🚀 🚨 After deployment, reports of misbehavior start trickling in What happened?? How could you have caught it?? 🤔 @icmlconf 2026 Spotlight! 🧵
Show more