๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    50.3K ํŒฌ
๐ŸŽ‰ Meet vLLM-Omni v0.22.0, a major upgrade for omnimodal world models and production-grade multimodal serving. ๐ŸŒ Day-0 @NVIDIAAI Cosmos 3 world models: text, image, audio, video, and action, in and out. ๐Ÿค– Robot serving: DreamZero + OpenPI realtime API. ๐ŸŽ™๏ธ Production TTS: Qwen3-TTS, Qwen3-Omni, VoxCPM2 and more. ๐ŸŽจ Faster image/video/diffusion: Wan 2.2, HunyuanVideo 1.5, LTX-2.3. โšก Broader quantization (FP8/INT8, MXFP4/MXFP8, W4A16, ModelOpt) and hardware coverage. 339 commits, 124 contributors, 52 of them new. Thank you all. ๐Ÿ™Œ ๐Ÿ”—
๋” ๋ณด๊ธฐ