It was one of the coolest, most thrilling tech adventures I’ve ever done. I got a wild buzz.
Qwen3.8 Flash Next just served a 262k token prompt with an image on my 4x3090. Decode at that depth: 67 tok/s. At 260K its 66.3, so the curve is basically FLAT.
Almost every number you see for this model is 20-21 tok/s on a single 24GB card with experts offloaded to RAM on llamacpp. I went the other way: vLLM, all experts on GPU, FP8 KV. Full recipe below in my github, one command, raw runs included.
🧵A deep-dive technical thread. Let's go!
顯示更多