Got access to the GLM 5.3 FlashX.
It's the same weights as 5.3 Flash, just served at 100 to 200 tok/s, for 2.5x the quota.
So it's not smarter and you're paying 2.5x for speed alone... dont you think 100-200 tok/s should be the standard for flash on the cloud? why a premium!?
Got temporary access to a Mac Studio M5 Ultra. Barely optimized, Qwen 3.8 Flash Next runs at 85.9 tok/s and 6085 tok/s prefill on long prompts.
A 2x DGX Spark recipe gets 52.1 tok/s single stream and about 2960 tok/s prefill on 16k–64k prompts.
Already ahead of 2x Spark, but I was expecting a much much better baseline…