註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Lucas Beyer (bl16)
@giffmana
Researcher (now: Meta. ex: OpenAI, DeepMind, Brain, RWTH Aachen), Gamer, Hacker, Belgian. Anon feedback: ✗DMs → email
加入 December 2013
648 正在關注    152.8K 粉絲
So they say "three things we scaled". Let me translate: 1. "compute" - yep, that's compute. 2. "environments and harnesses" - actually, also compute. 3. "and grader compute" - you guessed it, that's also compute. joke aside, pretty cool to see their public live dashboard, including "cost so far" (!) (due to recent events/discussions: no, this QT is not sponsored. I don't do sponsor stuff, I'm here for the fun.)
顯示更多
Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run:
顯示更多
0
16
490
17
轉發到社區