Register and share your invite link to earn from video plays and referrals.

Search results for A9
A9 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including A9
Qwen3.8-2.4T-A95B, compressed two ways at once: 25% of the experts pruned with REAP, and the rest quantized to NVFP4. Even with a quarter of the experts gone and 4-bit weights, GPQA Diamond holds at 91.5 vs 92.6 for the full-precision base. ~99% recovery. Serve on @vllm_project:
Show more
Qwen3.8-2.4T-A95B, a 2.4-trillion parameter MoE, quantized with no accuracy loss. MoE layers to NVFP4, attention layers to FP8 block, via LLM Compressor. On GPQA Diamond: 93.1 quantized vs 92.6 for the full-precision base. Full recovery. Serve on vLLM:
Show more
Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context window, and strong coding & agent capabilities. Start building:
Show more