注册并分享邀请链接,可获得视频播放与邀请奖励。

Alexander Doria
@Dorialexander
building open ai infrastructure @pleiasfr — χαλεπὰ τὰ καλά
加入 April 2011
4.1K 正在关注    24.5K 粉丝
Since I went into this release: *It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo. *Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm. *Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.
显示更多