登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Francesco Bertolotti
@f14bertolotti
AI Researcher
参加 October 2021
141 フォロー中    1.9K ファン
Stellar performance from a 3B model. These results were achieved primarily through post-training refinements on Qwen2.5-Coder. The paper doesn't provide many details, but it appears they distill from RL ckpts and then do a final RL-based instruct RL. 🔗
もっと見る