註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
加入 May 2018
2.1K 正在關注    12.6K 粉絲
Qwen3.8-2.4T-A95B, a 2.4-trillion parameter MoE, quantized with no accuracy loss. MoE layers to NVFP4, attention layers to FP8 block, via LLM Compressor. On GPQA Diamond: 93.1 quantized vs 92.6 for the full-precision base. Full recovery. Serve on vLLM:
顯示更多
0
1
152
12
轉發到社區