註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    46.7K 粉絲
🚀 Qwen3.6-27B-NVFP4 is inference ready with vLLM on NVIDIA Blackwell GPUs. This checkpoint is optimized for Blackwell and reduces GPU memory requirements by ~2.5x for local AI with open-source models. 🧠 27B params, Hybrid Attention 📊 NVFP4 evals: 86.3 on MMLU Pro, 85.5 on GPQA Diamond 🛠️ Exclusively supported on vLLM as the runtime engine Get started from the Hugging Face checkpoint:
顯示更多
Fast, efficient local AI with open-source models just got easier. Qwen3.6-27B-NVFP4 is now on @HuggingFace! It's optimized for NVIDIA Blackwell GPUs & inference ready with @vllm_project. The checkpoint reduces GPU memory requirements by approximately 2.5x for powerful 27B-parameter inference on your own hardware.
顯示更多
0
15
445
40
轉發到社區