DeepSeek 4.1 Flash full precision + DSpark @ 200TPS on 4 Max-Qs- with just 64GB system RAM!
Offloading the 200GB Engram (i.e. a hash table) to NVMe seems fine rather than keeping it all in RAM, my wallet is thankful ๐
More juice to squeeze though, vLLM recipe to come... ๐