this is the nightmare scenario we've been expecting. a glm 5.3 model with its safeguards removed completely independently, capable of unrestricted offensive cyber exploits. some thoughts:
- these guys downloaded glm 5.3's weights, ran a bunch of contrasting prompts to detect where the refusal safeguards were then used a technique known as orthagonalization to remove them.
- the resulting model accepts your requests 85% of the time. the "zero-refusal" version just spouts gibberish so doesn't actually work as well.
- this was only possible because the weights were open and available. this is essentially what dario was concerned abt.
- this will spark a secondary "grey market" for companies that seek to offer you the same model but effectively jail-broken.
with the hugging face exploit and anthropic + openai pausing training, we need to do everything to stop stuff like this becoming a problem.