This should not be possible.
Qwen3.8-27B just hit 70 tokens per second on a MacBook Pro M5 Max.
Same quality.
Up to 4.6ร faster than normal decoding.
A frontier-level open model running this fast on a laptop.
Local AI just became actually usable. ๐
DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro.
โก Up to 4.6ร the speed of autoregressive decoding, with the same output.
This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!