注册并分享邀请链接,可获得视频播放与邀请奖励。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在关注    50.2K 粉丝
vLLM integrated PyNvVideoCodec to offload video decoding from CPU to GPU's NVDEC. Result: 2x+ throughput at 8×H100, CPU bottleneck gone! This is huge for video captioning at scale (AV training, metadata), and ships with CUDA vLLM releases. 🔗
显示更多
0
4
88
15
转发到社区