๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Alex Ker ๐Ÿ”ญ
@thealexker
code+words @baseten | investing in frontiers & sharing my curiosities | prev @bloombergbeta @stanfordhai @neurable.
๊ฐ€์ž… May 2018
1.3K ํŒ”๋กœ์ž‰ ์ค‘    13.3K ํŒฌ
most people forget there are two vectors to optimize for to reduce model cost: 1) reducing the input/output costs, increasing cache hit rates, batching etc 2) packing more intelligence per token (fewer tokens for same task) inference cost = price per token ร— tokens per task optimizations around the second is underrated and something weโ€™ll only see more of. compression and concision is intelligence.
๋” ๋ณด๊ธฐ