Register and share your invite link to earn from video plays and referrals.

Edward Hu
@edwardjhu
Head of AI Modeling @ Mercor Ph.D. with Yoshua Bengio | ex-OpenAI
41 Following    9.2K Followers
How does one RL post-train a 397B model for long-horizon knowledge work? 👩‍💼 We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research. Full blog: Source code:
Show more
0
17
785
116
Forward to community
I am joining Mercor to lead model training & research. We envision a world with an abundance of intelligence. Great data is increasingly the bottleneck for frontier models to tackle economically valuable work. We believe great data can is best produced in connection with great model training. We commit to doing great modeling research out in the open, starting with sharing our 397B RL training run to hillclimb APEX-Agents, our flagship knowledge work benchmark:
Show more