註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    46.7K 粉絲
GLM-5.2 in NVFP4 is ready to serve in vLLM 🚀 @NVIDIAAI's official NVFP4 checkpoint of GLM-5.2 on Blackwell cuts the memory footprint vs FP8 while matching its accuracy across reasoning, coding, and long-context benchmarks. Serve it today with: vllm serve nvidia/GLM-5.2-NVFP4 🤗
顯示更多
0
9
319
25
轉發到社區