註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Thinking Machines
@thinkymachines
Thinking, beeping, and booping. @tinkerapi
加入 February 2025
1 正在關注    180.9K 粉絲
Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other approaches for a fraction of the cost.
顯示更多
0
59
2.8K
403
轉發到社區