註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Alex Ker 🔭
@thealexker
code+words @baseten | investing in frontiers & sharing my curiosities | prev @bloombergbeta @stanfordhai @neurable.
加入 May 2018
1.3K 正在關注    13.3K 粉絲
most people forget there are two vectors to optimize for to reduce model cost: 1) reducing the input/output costs, increasing cache hit rates, batching etc 2) packing more intelligence per token (fewer tokens for same task) inference cost = price per token × tokens per task optimizations around the second is underrated and something we’ll only see more of. compression and concision is intelligence.
顯示更多