Think your model is safe? Just ask it again and again!
Language models are probabilistic. We find that just repeatedly asking the same query, with no jailbreaks, can yield harmful outputs.
If your LLM is queried millions of times a day, these rare misbehaviors are inevitable!
显示更多