所谓的 API 蒸馏 因为必然是off-policy 的,主要是用于大规模合成 SFT 指令集用于RL前的冷启动阶段。
有帮助,但是只靠SFT无法制造出第一梯队的前沿模型。每家模型必须自建大规模的 RL 基础设施。
但是合成 SFT 指令集这件事上,我相信必然会利用某些外部强模型生成一些。
This paper shows that knowledge transfer is hard even under perfect conditions and model access, yet it is randomly used as an argument against what i was saying about distillation. Actually it is a good read to understand why API distillation can't work, and you can just fine tune for specific style/patterns or very narrow domain knowledge *at best*.
顯示更多