Congrats to
@Zai_org on GLM-5.3-Flash! Day 0 support in Nativ 🎉
320B params, 18B active · 1M native context · fully on-device.
On an M3 Ultra (512GB) with Nativ v0.3.4, standard 4-bit MLX:
⚡ Up to 505 tok/s prefill · 32 tok/s decode
🚀 Batch 4 lifts decode to 70.9 tok/s
💾 Peak under 380GB, fully resident, no offload
📏 Benchmarked to 128K context
A 320B MoE running entirely on your Mac.
Try Nativ 👇