DeepSeek-V4-Flash-0731 (all 284B parameters) running on a single DGX Spark.
⚡ ~17 tok/s generation, 40 tok/s prefill
📦 One 80GB GGUF file, mixed IQ2_XXS/Q8 imatrix quant
🧠 256-expert MoE, 6 routed per token
🆓 MIT licensed
frontier model (kinda) on your desk 😎
🤗