Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far.
Fastest right now: Ollama's MLX engine (DFlash included) at ~29 tok/s.
Tuned llama.cpp: ~21. Raw mlx-vlm: ~10, not optimized yet.
Numbers below if you're setting it up 👇