๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    50.4K ํŒฌ
โšก New RTX local-agent optimizations include a vLLM speedup on Blackwell. @NVIDIARTXSpark reports: ๐Ÿ› ๏ธ 1.2x vLLM performance on RTX PRO 6000 Blackwell โšก Up to 1.4x on a two-system DGX Spark cluster Great to see the local serving path getting faster across RTX and DGX Spark.
๋” ๋ณด๊ธฐ