注册并分享邀请链接,可获得视频播放与邀请奖励。

日常焦虑帝
@gpuhell
攻城狮/业余投机/右侧交易/游戏开发/Haskell/Rust/C++/Unity3D/C#/对java有偏见/RL/智商欠费/浅尝辄止故平庸/反乌托邦/竹林中/地狱变/手撕菠萝蜜/胸口碎榴莲/单机推特中/乐视一生黑(乐视已阵亡)/华为一生黑(迟早会阵亡 )/中华跪族/器材党/预防式B台支
加入 June 2012
2.5K 正在关注    1.1K 粉丝
The previously missing data from MiMo’s RL training has now been added.
MIMO’s RL training has stopped at step 30. We can observe the following: The Pro model achieved a DeepSWE score of 72.57, but this score was reported at step 28; data for steps 29 and 30 are missing. The number of active environments for Pro began increasing at step 23, then dropped rapidly after step 26. The duration of each training step also rose sharply.
显示更多