MIMO’s RL training has stopped at step 30. We can observe the following:
The Pro model achieved a DeepSWE score of 72.57, but this score was reported at step 28; data for steps 29 and 30 are missing.
The number of active environments for Pro began increasing at step 23, then dropped rapidly after step 26. The duration of each training step also rose sharply.