注册并分享邀请链接,可获得视频播放与邀请奖励。

Modular
@Modular
Building AI’s unified compute layer. We are hiring → 🚀
加入 January 2022
2 正在关注    24.3K 粉丝
Optimizing large scale inference systems is what we do, so we decided to write down what we know. Our LLM Inference Handbook is a free reference covering TTFT, TPOT, goodput, continuous batching, chunked prefill, prefix caching, KV cache math, prefill-decode disaggregation, quantization, and more. It includes 20+ interactive visualizations, is updated continuously, and is open to PRs.
显示更多
0
2
285
47
转发到社区