It is very easy to hide all bugs of Lean in a correct looking proof, so it’s vital that there are no holes!
We can just run models on some very hard problems and detect the holes early, but from an alignment point of view we have to make sure models don’t use the holes when asked to formalize.
Here is our postmortem describing new Lean bugs found by OpenAI internal models. They are all fixed in Lean v4.33.1 Many thanks to Daniel Selsam from OpenAI for all the help.