the weird part of this MiMo graph is that the judge and probe disagree 60% of the time.
and for pro, disagreement goes up during training.
would love to know how much of that is harder-to-judge behavior vs the model getting better at fooling the judge.