注册并分享邀请链接,可获得视频播放与邀请奖励。

KVCache.AI
@KVCache_AI
Hi, this is official account. We build systems for efficient LLM serving, including KTransformers, Mooncake and AgentENV.
加入 August 2018
109 正在关注    1.1K 粉丝
Agentic inference is making KV cache management and data movement increasingly central to the serving stack. The new AgentX / InferenceXv3 analysis from @SemiAnalysis_ captures this shift well. Glad to see Mooncake highlighted for external KV caching and disaggregated P/D data transfer, including support for hybrid-attention cache management and preserving speculative decoding state. Special thanks to AMD and SemiAnalysis for providing the CI resources and technical support that made Mooncake ROCm wheel packaging possible. We’re excited to ship it in our next release. Read more:
显示更多
0
1
41
10
转发到社区