注册并分享邀请链接,可获得视频播放与邀请奖励。

Nathan Lambert
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote
加入 December 2014
943 正在关注    102K 粉丝
One of the coolest at-scale RL resources made public yet! You love to see it.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
显示更多
0
13
621
27
转发到社区