๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Edward Hu
@edwardjhu
Head of AI Modeling @ Mercor Ph.D. with Yoshua Bengio | ex-OpenAI
๊ฐ€์ž… December 2019
41 ํŒ”๋กœ์ž‰ ์ค‘    9.2K ํŒฌ
How does one RL post-train a 397B model for long-horizon knowledge work? ๐Ÿ‘ฉโ€๐Ÿ’ผ We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script๐Ÿš€ This is the first of many works from Mercor Research on open model training research. Full blog: Source code:
๋” ๋ณด๊ธฐ