Theorem co-founders
@rajashree +
@diagram_chaser explain why formal verification could break cybersecurity’s endless whack-a-mole and stop reward hacking during AI training:
Rajashree Agrawal: "Formal verification is asymmetric security. Currently you have this whack-a-mole problem. As the attackers get better, the defenders have to catch up. You keep doing this forever."
"Once you prove this particular property holds of your program, you don't need to check that again. This feels like the asymmetric approach that you need if you're going to avoid this cyber apocalypse."
"If you wanna move beyond one-off interactions, build a very complex world, lots of automations, you're going to have structured things coming out of models which looks like software. Being able to reason about software seems like a good hammer to have."
Jason Gross: "If you can verify the RL environments and the graders that you're running, then the models won't have any reward hacks during training, and so you can potentially train them to not be as reward hacky in general."
@theoremlabs