가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Awni Hannun
@awnihannun
ow knee
가입 January 2011
358 팔로잉 중    45K 팬
Qwen 27B dense at 105 tok/s output on an m5 max is pretty bonkers. Breaking down the memory wall one brick at a time.
We’re releasing speculative decoding in our inference engine Uzu, starting with Qwen3.6-27B. On Apple M5 Max with 128 GB of unified memory, our Mirai-M stack reaches 105 output tokens/sec entirely on-device - 2.9× faster than the fastest MLX speculative-decoding implementation we benchmarked. The result comes from full-stack co-design: DFlash + our Weaver model, tree-based speculative decoding, Mirai quantization, our verification algorithm, and Metal kernels optimized for Apple silicon. Mirai builds the full stack for frontier on-device intelligence. 105 t/s is the evidence, full-stack integration is the moat. But tokens/sec are not actually the metric we ultimately care about. More on that soon. Full benchmarks + methodology:
더 보기