註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jesse Zhang
@thejessezhang
加入 October 2021
203 正在關注    102.1K 粉絲
Big +1 to this. We were a bit skeptical of this when we were first starting out, but once the task is defined, it's strictly better to move to smaller models. Better latency, better cost, AND you can tune it for the specific task. Great reason for frontier labs to open-source small models since they don't compete with the big models on the same use cases.
顯示更多
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task.
顯示更多
0
10
122
8
轉發到社區