註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Edward Hu
@edwardjhu
Head of AI Modeling @ Mercor Ph.D. with Yoshua Bengio | ex-OpenAI
加入 December 2019
42 正在關注    9.3K 粉絲
How does one RL post-train a 397B model for long-horizon knowledge work? 👩‍💼 We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research. Full blog: Source code:
顯示更多
0
16
791
117
轉發到社區