็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

Cruz
@ViC305
God-fearing husband & father. AI Engineer. Fine-tuning + day-zero GGUFs + benchmarks โ€ข Local LLMs โ€ข honest tok/s โ€ข fitting big models on small GPUs
ๅ‚ๅŠ  June 2009
828 ใƒ•ใ‚ฉใƒญใƒผไธญ    1.1K ใƒ•ใ‚กใƒณ
๐Ÿณ DeepSeek-V4-Flash-Vision-Exp is 305B, and I packed it into a ~95 GiB mixed-precision EXL3 release that now SERVES on one DGX Spark. Receipts + Recipe below. ๐Ÿš€ K2 across the routed-expert stack. K3 only where my layer sensitivity scan ranked the extra precision highest. MTP/DSpark drafter tensors preserved from source. ๐— ๐—œ๐—ซ๐—˜๐—— ๐—ž, ๐—ก๐—ข๐—ง ๐—™๐—Ÿ๐—”๐—ง ๐Ÿฎ-๐—•๐—œ๐—ง The source checkpoint is already mixed format: MXFP4 routed experts FP8 attention MXFP4 MTP/DSpark drafter I dequantized each routed-expert tensor to BF16 before trellis encoding, then rebuilt the expert bank as: 43 MoE layers on a K2 base K3 on layers 3, 13, 21, 22, 28, 41 attention / shared experts / embeddings kept at high precision (BF16) 48 output shards, ~95 GiB total Those six layers came from a complete K2 vs K3 proxy-error scan, not guesses. The ranking was nearly flat across all 43 layers, so I am not pretending these are six magical outliers. The real question is whether the extra bits there produce a measurable fidelity gain over uniform K2 that comparison is still open. ๐—ฆ๐—ง๐—œ๐—Ÿ๐—Ÿ ๐—ข๐—ฃ๐—˜๐—ก prefill/decode numbers + max ctx PPL/KLD vs uniform K2 MTP/DSpark speculative serving CUDA-graph decode numbers sixcat battery (first full run is in; publishing after a chat-template fix that currently understates tool use) Model: Serve it (one Spark, open source):
ใ‚‚ใฃใจ่ฆ‹ใ‚‹