Register and share your invite link to earn from video plays and referrals.

Alex Ker 🔭
@thealexker
code+words @baseten | investing in frontiers & sharing my curiosities | prev @bloombergbeta @stanfordhai @neurable.
Joined May 2018
1.3K Following    13.3K Followers
most people forget there are two vectors to optimize for to reduce model cost: 1) reducing the input/output costs, increasing cache hit rates, batching etc 2) packing more intelligence per token (fewer tokens for same task) inference cost = price per token × tokens per task optimizations around the second is underrated and something we’ll only see more of. compression and concision is intelligence.
Show more