๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

KVCache.AI
@KVCache_AI
Hi, this is official account. We build systems for efficient LLM serving, including KTransformers, Mooncake and AgentENV.
๊ฐ€์ž… August 2018
109 ํŒ”๋กœ์ž‰ ์ค‘    1.1K ํŒฌ
๐Ÿš€ Over one trillion tokens per day, powered by Mooncake and SGLang. @ApproachingAI shares the system architecture and engineering lessons behind its token factory, running leading trillion-parameter models with SGLang HiCache + Mooncake Store. A deep dive into: ๐Ÿ—๏ธ Production architecture and system design trade-offs โšก Performance optimizations for trillion-token-scale workloads ๐Ÿ›ก๏ธ Failure isolation, fault tolerance, and recovery Real production lessons on scaling KV cache reuse while meeting strict inference SLOs. Read more:
๋” ๋ณด๊ธฐ