How does one RL post-train a 397B model for long-horizon knowledge work? 👩💼
We share every step we took to bring Qwen 3.5 397B from 16.1% Pass
@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research.
Full blog:
Source code: