註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

himanshu
@himanshustwts
cofounder @PhyseraAI • host @groundzero_twt • DMs open!
加入 April 2022
3.7K 正在關注    29.1K 粉絲
BREAKING from Chinese frontier: They are live streaming big RL runs now. God hail open source like wdym you can learn so much from here just by metric visualizations.
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
顯示更多
0
8
980
61
轉發到社區