가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Modular
@Modular
Building AI’s unified compute layer. We are hiring → 🚀
가입 January 2022
2 팔로잉 중    24.3K 팬
Optimizing large scale inference systems is what we do, so we decided to write down what we know. Our LLM Inference Handbook is a free reference covering TTFT, TPOT, goodput, continuous batching, chunked prefill, prefix caching, KV cache math, prefill-decode disaggregation, quantization, and more. It includes 20+ interactive visualizations, is updated continuously, and is open to PRs.
더 보기