๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
280 ํŒ”๋กœ์ž‰ ์ค‘    413 ํŒฌ
๐Ÿ’พ TL;DR: reuse a KV cache computed on another node's GPU straight through CXL memory, up to 36.6x faster. Seagate proposes a Kubernetes-native way to share CXL memory across the cluster. Title: Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving URL: ๐Ÿ“Œ Highlights ๐Ÿงฉ A Kubernetes DRA driver turns CXL regions into schedulable cluster resources ๐Ÿ—‚ No Redis/etcd needed, KV slots are tracked inside the region itself ๐Ÿš€ Cross-node reuse hits 36.6x speedup over recompute at 32K-token prefixes ๐ŸŽฏ Hit rates reach 95.4-99.5% โš–๏ธ Cross-node vs same-node latency differs by only 1-4% ๐Ÿข Still 1.15-1.70x slower than an on-GPU VRAM prefix cache hit ๐Ÿ”ง Measured 5.8 GB/s out of a 27 GB/s fabric ceiling, room to grow with zero-copy DMA For serving that needs to reuse long prefixes across GPU nodes, this feels like a genuinely practical option. #LLMInference# #Kubernetes#
๋” ๋ณด๊ธฐ