註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

KVCache.AI
@KVCache_AI
Hi, this is official account. We build systems for efficient LLM serving, including KTransformers, Mooncake and AgentENV.
加入 August 2018
109 正在關注    1.1K 粉絲
Agentic inference is making KV cache management and data movement increasingly central to the serving stack. The new AgentX / InferenceXv3 analysis from @SemiAnalysis_ captures this shift well. Glad to see Mooncake highlighted for external KV caching and disaggregated P/D data transfer, including support for hybrid-attention cache management and preserving speculative decoding state. Special thanks to AMD and SemiAnalysis for providing the CI resources and technical support that made Mooncake ROCm wheel packaging possible. We’re excited to ship it in our next release. Read more:
顯示更多
0
1
41
10
轉發到社區