Since I went into this release:
*It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo.
*Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm.
*Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.