登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Nathan Lambert
@natolambert
Open model research @ something new. Prev. co-led Olmo at Ai2. Writes @interconnectsai, wrote
参加 December 2014
945 フォロー中    102.3K ファン
The biggest disagreement JSD and I had in this podcast was on the impact of distillation. In our research for Interconnects, @xeophon and I agree that there's no hard evidence that distillation is a massive impact for the Chinese labs. At the same time, the gossip mill in SF has been doubling down on the Chinese labs getting massive gains from distilltion. The argument where distillation is a huge impact is something along the lines of distilled traces go into mid training and make RL work far more easily. My argument, that distillation is a 1-2 month pull ahead in capabilities closer to the frontier, is that Chinese labs already have sufficiently strong models, where this mid-training setup can be done on their own models, and most of the capabilities gains are from scaling RL environment training, which doesn't link cleanly to the distillation data pathway. We feel like the recent RL dashboard from @XiaomiMiMo supports this claim. A lot of it comes down to a gut call on if you can believe it that the Chinese labs are really great at building LLMs, maybe even better than focusing than OpenAI/Ant etc, as their current ambitions are a bit narrower (catching up), rather than transformative products/inventions
もっと見る
New podcast with @datagenproc of @EpochAIResearch digging into the open questions determining the future of frontier AI! We cover: 00:00 Predictions for RSI 18:15 The role of robotics in an AI acceleration 24:20 How far behind are Chinese models? 27:39 Does distillation explain the gap? 40:58 What Chinese job postings reveal about their labs 48:13 Are open or closed models safer? 58:10 How Epoch AI ticks 1:00:55 What a frontier post-training recipe looks like He's one of the people who gives the best feedback on my writing, so I was stoked to have him on.
もっと見る