注册并分享邀请链接,可获得视频播放与邀请奖励。

Jesse Zhang
@thejessezhang
加入 October 2021
203 正在关注    102.1K 粉丝
Big +1 to this. We were a bit skeptical of this when we were first starting out, but once the task is defined, it's strictly better to move to smaller models. Better latency, better cost, AND you can tune it for the specific task. Great reason for frontier labs to open-source small models since they don't compete with the big models on the same use cases.
显示更多
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
显示更多
0
10
122
8
转发到社区