Unfortunately, we had to remove this entry: the authors ran this as an autoformalization task with access to source papers. Our pipeline run the models without such access.
80% is still impressive for open models on this challenging task, just not comparable to the other scores