註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Haseeb Qureshi >|<
@hosseeb
Managing partner @DragonflyVC. Let's think step by step.
加入 December 2010
1.2K 正在關注    149.1K 粉絲
Excellent primer on the economics of token caching, and why fancy token model routers aren't as economical as you'd think. As workloads lean more agentic (way more tokens in than out vs chat), caching dominates as the most important input into token costs.
顯示更多
Wrote about the incredibly important KV cache: - Why are people spending so much on agents - How is inference so profitable all of a sudden (+ why it won't last) - Why routing doesn't work P.S. they also constrain the design space for coding agents A LOT
顯示更多
0
11
33
2
轉發到社區