註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
加入 May 2018
2.1K 正在關注    12.3K 粉絲
Qwen3.8-2.4T-A95B, compressed two ways at once: 25% of the experts pruned with REAP, and the rest quantized to NVFP4. Even with a quarter of the experts gone and 4-bit weights, GPQA Diamond holds at 91.5 vs 92.6 for the full-precision base. ~99% recovery. Serve on @vllm_project:
顯示更多
0
3
186
22
轉發到社區