註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Awni Hannun
@awnihannun
ow knee
加入 January 2011
358 正在關注    45K 粉絲
Qwen 27B dense at 105 tok/s output on an m5 max is pretty bonkers. Breaking down the memory wall one brick at a time.
We’re releasing speculative decoding in our inference engine Uzu, starting with Qwen3.6-27B. On Apple M5 Max with 128 GB of unified memory, our Mirai-M stack reaches 105 output tokens/sec entirely on-device - 2.9× faster than the fastest MLX speculative-decoding implementation we benchmarked. The result comes from full-stack co-design: DFlash + our Weaver model, tree-based speculative decoding, Mirai quantization, our verification algorithm, and Metal kernels optimized for Apple silicon. Mirai builds the full stack for frontier on-device intelligence. 105 t/s is the evidence, full-stack integration is the moat. But tokens/sec are not actually the metric we ultimately care about. More on that soon. Full benchmarks + methodology:
顯示更多
0
22
345
20
轉發到社區