登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

wh
@nrehiew_
eng primarily, ml mostly, research previously
参加 October 2023
104 フォロー中    18.5K ファン
Ton of detail here, even including data and harness composition, wow
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
もっと見る