注册并分享邀请链接,可获得视频播放与邀请奖励。

Matt White
@matthew_d_white
Open Intelligence | Ex-Global CTO of AI, Linux Foundation | Ex-Exec Dir/CTO, PyTorch Foundation | UC Berkeley | Columbia University | Researcher & Educator
加入 March 2014
1.2K 正在关注    2.3K 粉丝
We absolutely need more labs to provide this level of transparency into all stages of training. All American and Chinese labs should follow suit. Well done @_LuoFuli and the Shaomi team.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
显示更多