註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    50.2K 粉絲
vLLM integrated PyNvVideoCodec to offload video decoding from CPU to GPU's NVDEC. Result: 2x+ throughput at 8×H100, CPU bottleneck gone! This is huge for video captioning at scale (AV training, metadata), and ships with CUDA vLLM releases. 🔗
顯示更多
0
4
88
15
轉發到社區