가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

SemiAnalysis
@SemiAnalysis_
가입 January 2024
35 팔로잉 중    167.9K 팬
Most production inference services run on Kubernetes, where declarative configuration helps coordinate distributed services. Running one inference server in a Slurm job is easy; running a production-style topology across multiple containers and nodes is not. NVIDIA’s open-source srt-slurm ( provides a YAML-based orchestration layer for these deployments. It coordinates sbatch and srun, GPU placement, networking, readiness checks, and cleanup for systems involving prefill/decode disaggregation, multiple workers, Dynamo frontends, KV-cache-aware routers, and KV-cache offloading services. It’s useful for quickly deploying representative inference systems for testing or benchmarking on GPU clusters that already run Slurm. Great work to @0xishand and the rest of the @NVIDIAAI team on this!
더 보기