Register and share your invite link to earn from video plays and referrals.

Francesco Bertolotti
@f14bertolotti
AI Researcher
141 Following    1.9K Followers
Stellar performance from a 3B model. These results were achieved primarily through post-training refinements on Qwen2.5-Coder. The paper doesn't provide many details, but it appears they distill from RL ckpts and then do a final RL-based instruct RL. 🔗
Show more